Bisulfite sequencing in DNA methylation analysis
Bisulfite sequencing (BS-seq) is a powerful technique that has revolutionized the study of DNA methylation, a key epigenetic modification where methyl groups are added to cytosine bases, typically at CpG dinucleotides.
By treating DNA with sodium bisulfite, unmethylated cytosines are converted to uracil, while methylated cytosines remain unchanged, allowing precise mapping of methylation patterns at single-nucleotide resolution. Comparing bisulfite-treated and untreated samples reveals intricate methylation landscapes across various organisms, cell types, and tissues. This method provides critical insights into gene regulation, cellular differentiation, and disease processes, such as cancer.
DNA methylation
DNA methylation, an epigenetic modification adding a methyl group to cytosine (forming 5-methylcytosine, typically at CpG sites), regulates transcription, X chromosome inactivation, transposon silencing, genomic imprinting, embryonic development, and chromatin structure. It represses gene expression at promoter regions and CpG islands by blocking transcription factors or attracting repressive proteins, but can promote expression in gene bodies. Altered methylation patterns are linked to diseases like cancer, making it a key epigenetic mark for gene regulation.
DNA methylation is catalyzed by DNA methyltransferases (DNMTs) in mammals. The DNMT family is comprised of enzymes with distinct roles in DNA methylation. For example, DNMT1 is primarily responsible for maintaining existing methylation patterns during DNA replication. DNMT3a and DNMT3b are involved in de novo methylation, establishing new methylation patterns during development or in response to environmental cues.
Analysis of high-throughput DNA methylation has become a regular phenomenon in laboratories worldwide. These methods include restriction enzyme-based approaches, bisulfite conversion and affinity-based approaches. Bisulfite sequencing is widely regarded as the gold standard for DNA methylation analysis due to its high resolution and accuracy. This method provides a single-base resolution, making it a powerful tool for studying epigenetic modifications across the genome.
How does bisulfite sequencing work?
The key to bisulfite sequencing is sodium bisulfite, which converts unmethylated cytosines to uracil via deamination. Methylated cytosines are protected from the action of sodium bisulfite and so remain unconverted during the process. Conversion is followed by amplification, where the uracils generated by the conversion process are amplified as thymines. The samples are then sequenced to analyze methylation patterns at single-base resolution. Each stage is described in more detail in the next section.
Figure. Steps in bisulfite treatment, including preservation of methylated cytosines
Sample preparation and DNA extraction
The first stage of bisulfite sequencing is DNA extraction, which involves isolating pure, high-quality DNA from a biological sample to ensure accurate and reliable sequencing results. Ensuring the DNA is free from contaminants, such as proteins or RNA, will improve the efficiency and accuracy of the sequencing process. Genomic DNA can be isolated from various sources using commercial DNA extraction kits.
Modifications to bisulfite protocols can allow the analysis of low-quantity DNA when necessary. For example, while fresh tissue samples work well, samples from paraffin-embedded tissues may yield poor results due to DNA degradation. It has been reported that the sequence libraries were 10% lower in formalin-fixed, paraffin-embedded (FFPE) tissue than in fresh-frozen tissue1, 2.
To address the issues with FFPE samples, a variation of reduced representation bisulfite sequencing (RRBS) protocol has been developed specifically for FFPE samples. These protocols incorporate steps such as end-polishing, optimal buffer selection, and performing all enzyme reactions in the same tube to enhance the efficiency of the process1, 2.
Bisulfite treatment
Bisulfite sequencing uses sodium bisulfite to convert unmethylated cytosines to uracils via hydrolytic deamination, preserving methylated cytosines. The process can cause DNA fragmentation and requires sufficient DNA input for optimal recovery. After Bisulfite conversion, the sample will need to undergo desulphonation and clean-up before moving to the next stage. Commercial kits like bisulfite conversion kit – whole cell streamline the procedure, though variations may exist in incubation times and DNA input requirements.
PCR amplification
After conversion, the samples are amplified by polymerase chain reaction (PCR). The 5-methylcytosines are amplified as cytosines, whereas the uracils generated by the conversion process are amplified as thymines. This means that bisulfite-treated DNA is usually AT-rich and has a low GC composition. Therefore, non-specific PCR amplification is relatively common with bisulfite-treated DNA, and it is strongly recommended to perform the PCR using high-fidelity “hot start” polymerases to reduce the error rate at this critical stage3.
Unlike normal PCR, bisulfite PCR requires longer primers, typically ranging from 26 to 30 bases, and shorter amplicons, usually between 150 and 300 bp. Ideally, primers should avoid CpG sites, but, if necessary, they should be positioned at the 5'-end with a mixed base at the cytosine position. Successful amplification generally requires 35 to 40 cycles. Annealing temperatures between 55°C and 60°C are typically effective, and running an annealing temperature gradient with each new primer set helps optimize target-specific amplification.
Sequencing
Sanger sequencing is a reliable method for analyzing DNA methylation at specific CpG sites, offering qualitative methylation data, but it can be time-consuming and limited to small-scale analysis.
In contrast, next-generation sequencing (NGS) allows high-throughput, quantitative analysis of DNA methylation with greater resolution and multiplexing capability. However, it requires specialized equipment and bioinformatics expertise. Additionally, different library preparation methods can affect the performance of bisulfite-treated DNA in NGS. Optimizing these methods is crucial for achieving high coverage and precision.
Post-sequencing data generation
Post-sequencing data generation involves analyzing DNA methylation patterns by mapping reads, identifying methylated cytosines, and assessing coverage and quality. The type of post-sequencing analysis required depends on the specific bisulfite sequencing method used.
Types of bisulfite sequencing techniques
Bisulfite sequencing methods are tailored to specific needs, such as full-genome analysis, cost-effective partial coverage, or targeted examination of specific regions.
Whole genome bisulfite sequencing (WGBS)
Whole genome bisulfite sequencing combines bisulfite treatment with high-throughput next-generation sequencing (NGS), enabling researchers to explore DNA methylation patterns at single-nucleotide resolution across the entire genome. WGBS acts as a powerful tool for comprehensive epigenomic studies offering high resolution, making it essential for understanding gene regulation and disease mechanisms4.
Reduced representation bisulfite sequencing (RRBS)
Reduced representation bisulfite sequencing is a method used to analyze DNA methylation by selectively enriching certain regions of the genome. This method employs restriction enzymes like MspI in mammals to digest genomic DNA, targeting CpG-rich regions.
The regions containing CpGs are then enriched by size selection, often using gel purification. After methylated adaptor ligation, end repair and dA tailing, these fragments are then converted by bisulfite, amplified, and sequenced. It provides single-nucleotide resolution methylation data across a significant portion of the genome. By focusing on CpG-rich regions, RRBS reduces the amount of sequencing required and, therefore, lowers costs compared to whole-genome bisulfite sequencing5.
Targeted bisulfite sequencing (TBS)
Targeted bisulfite sequencing is a focused approach for analyzing DNA methylation in specific genomic regions, offering higher sequencing depth and coverage compared to WGBS. It is a cost-effective and highly targeted method for DNA methylation analysis, ideal for studies requiring detailed methylation analysis of specific genes or regions. Examples include validation studies, gene regulation research, and clinical methylation marker screening6.
Oxidative bisulfite sequencing (oxBS-Seq)
Oxidative bisulfite sequencing enables the absolute quantification of both 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) at single-base resolution, addressing the limitation of standard bisulfite sequencing.
In oxBS-seq, 5hmC is oxidized to 5-formylcytosine (5fC) using an oxidizing agent. The DNA is then treated with bisulfite, which converts 5fC to uracil, similar to non-methylated cytosine. By comparing BS-seq and oxBS-seq data sets, researchers can accurately identify 5hmC sites7.
Data production, quality control, and initial processing for NGS data
Bisulfite sequencing employs various methods for data control and processing tailored to specific needs. Quality control is maintained by ensuring accurate conversion of unmethylated cytosines, and by assessing read quality.
Initial processing includes trimming and filtering low-quality bases and adapter sequences from the reads. This is followed by calling methylation states using software designed for bisulfite-treated DNA, which generates a methylation report that includes the methylation status of each cytosine in the genome.
Quality control
Conversion efficiency: Assess the efficiency of bisulfite conversion by checking the proportion of unmethylated cytosines converted to uracils. High conversion efficiency is crucial for accurate methylation analysis. For example, conversion efficiency can be checked by PCR with the bisulfite-converted DNA regular non-bisulfite specific primers to amplify unconverted products.
Read quality: Evaluate the quality of sequencing reads using tools like FastQC to identify any issues with read length, base quality, or adapter contamination.
Coverage assessment: Ensure sufficient coverage of the target regions to provide reliable methylation data. Low coverage can lead to incomplete or biased results.
Quality checks: Perform quality checks of raw FASTQ files before and after trimming with FastQC.
Spiked-in controls: Use completely methylated or completely unmethylated “spiked in” controls in different libraries to assess the methylation status (0 or 100%), reflecting the data quality4, 8, 9, 10, 11.
Data normalization
Normalize the methylation data to account for any technical biases and ensure comparability across samples 4, 8, 9, 10, 11. Examples of normalization approaches include:
Read count normalization: Dividing the methylation read counts by the total number of sequenced reads (or a suitable fraction of the total reads) can provide a basic normalization approach.
Normalization by coverage: Adjusting methylation read counts based on the coverage depth of each region or CpG site helps account for variations in sequencing depth.
Statistical normalization: Various statistical methods, including quantile normalization, can be used to adjust for biases and differences in methylation levels between samples.
Normalization by internal standards: Including normalization domains or controls in the bisulfite sequencing protocol can allow for more precise and accurate normalization.
Bisulfite sequencing analysis and data interpretation
The analysis of DNA methylation sequencing data can be challenging for beginners due to the wide selection of tools and pipelines available, with the need for better guidelines to help researchers process, visualize, and analyze the data effectively, particularly when comparing conditions like cancer and normal tissues. Bioinformatics tools like short-read aligners are essential for analyzing bisulfite sequencing data, offering precise alignment, differential methylation analysis, and improved detection of differentially methylated sites.
Methylation pattern identification and differential methylation analysis
- Methylation pattern identification: Detects the distribution of DNA methylation across the genome, highlighting regions where methylation is present or absent.
- Differential methylation analysis: Compares methylation levels between different conditions or sample groups to identify changes associated with biological processes or diseases.
Visualization and interpretation of methylation results
Visualization and interpretation of methylation results involve analyzing methylation patterns through graphical representations, such as heatmaps or scatter plots, to identify significant variations and their biological implications.
Applications of bisulfite sequencing in research and medicine
Bisulfite sequencing is widely used in research and medicine to study DNA methylation, providing insights into gene regulation, disease mechanisms, and the development of diagnostic and therapeutic tools.
Cancer genomics
Commonly studied genes in cancer research include tumor protein p53 (TP53), breast cancer gene 1 (BRCA1), and Ras association domain family 1 isoform A (RASSF1A), where methylation alterations are linked to tumorigenesis and cancer prognosis. The potential reversibility of methyltransferase activity makes it an attractive target for therapeutic interventions, unlike genetic changes.
Bisulfite sequencing can detect early DNA methylation changes in circulating tumor DNA (ctDNA) from blood, offering a non-invasive and accurate method for cancer diagnosis12, 13.
Developmental biology
BS-seq helps in understanding methylation patterns in different developmental stages, such as during early embryogenesis, and can reveal how DNA methylation regulates the silencing of pluripotency genes and the activation of lineage-specific genes.
DNA methylation patterns change during development, with two main waves of demethylation and remethylation in humans. The first wave occurs in the germline, where global methylation is erased and sex-specific methylation patterns are established. The second wave happens after fertilization, involving the erasure of most methylation marks inherited from the gametes and the establishment of embryonic methylation patterns14, 15.
Neurobiology and disease research
BS-seq has also been used to investigate methylation changes associated with neurodegenerative diseases like Alzheimer's, Parkinson's, and Huntington's disease. By analyzing methylation patterns in brain tissues or cell-free DNA (cfDNA) from blood plasma, researchers hope to identify biomarkers for early diagnosis and track disease progression16.
Challenges and limitations of bisulfite sequencing
Challenges in studying DNA methylation changes include complex tissue composition, cell-type specificity, and uncharacterized genomic regions. Fortunately, emerging technologies are offering promising insights into these dynamics.
Sensitivity to DNA quality and potential degradation
BS-seq faces limitations such as DNA damage and incomplete conversion due to degradation. To address these issues, ultrafast bisulfite sequencing (UBS-seq) employs high temperatures and elevated bisulfite concentrations. This approach overcomes the long processing times of traditional bisulfite sequencing, improving efficiency by accelerating reactions and reducing DNA damage.
Technical challenges in bisulfite treatment and PCR biases
The PCR amplification step in bisulfite sequencing can introduce significant biases due to differential amplification efficiencies influenced by DNA methylation status. Non-methylated CpG sites are often amplified more efficiently than methylated ones, potentially leading to an underrepresentation of methylated regions and inaccurate methylation profiles.
Additionally, PCR amplification can exacerbate these issues through mispriming and the generation of high levels of duplicate reads. To mitigate these biases, amplification-free library preparation methods have been developed. Alternatives such as enzymatic methyl sequencing (EM-seq) and advanced techniques like terminal deoxynucleotidyl transferase (TdT)-assisted adenylate connector-mediated single-stranded DNA (TACS) ligation improve throughput and accuracy by enabling longer library fragment construction and reducing amplification artifacts.
Data complexity and interpretation difficulties
The complexity of large-scale bisulfite sequencing data, despite its potential for advancing genomic studies, poses challenges due to its difficult utilization. Sequencing errors, contaminants, variable read depths, and missing data further complicate interpretation and reduce reproducibility.
Certain omics platforms address this issue by integrating and simplifying access to diverse omics data, enabling efficient analysis and discovery of key genetic traits.
Developing standardized computational tools and optimized protocols for data filtering can enhance the accuracy and reproducibility of bisulfite sequencing analysis.
Best practices for performing bisulfite sequencing
Best practices are required to ensure accurate, reliable results and minimize biases in bisulfite sequencing data analysis.
Ensuring high-quality bisulfite treatment and sequencing
Optimizing bisulfite conversion conditions and using amplification-free library preparation methods, such as post-bisulfite adaptor tagging (PBAT) with TACS ligation, can reduce biases and improve sequencing quality.
Practical steps to prevent biases during bisulfite treatment
To prevent biases during bisulfite treatment:
Minimize PCR amplification: The methods that do not require PCR to reduce biases should be used. For the utilization of PCR, the right bisulfite conversion methods and polymerases need to be used for the reduction of errors.
Quality control: The quality of data needs to be checked before and after alignment, and tools that help in spotting biases can be used.
Library preparation protocols: The protocols that provide the coverage of CpG sites to be used and methods that create duplicates should be avoided. Library preparation kits like the post-bisulfite DNA library preparation kit (for Illumina®) can be a practical approach for preparing a library for all kinds of bisulfite sequencing methods like WGBS, oxBs-Seq and RRBS.
Recent advances and future directions in bisulfite sequencing
Advances in bisulfite sequencing aim to enhance its accuracy and scalability, paving the way for broader applications in epigenetic research and clinical diagnostics.
Single-cell bisulfite sequencing and its applications
Single-cell bisulfite sequencing (scBS) is a technique that allows for the profiling of DNA methylation patterns at single-base pair resolution in individual cells. By analyzing DNA from individual cells, scBS allows for the identification of cell-specific methylation patterns and the detection of variations between cells within a tissue or population.
Drop-BS is a droplet-based microfluidic technology that enables high-throughput single-cell bisulfite sequencing for DNA methylome profiling, allowing the analysis of up to 10,000 single cells in 2 days to reveal cell type heterogeneity in mixed cell lines and tissues. The approach was based on bisulfite conversion in droplets with DNA from a single cell with a unique barcode17.
Additionally, new techniques like bisulfite-converted randomly integrated fragments sequencing (BRIF-seq) and Smart-RRBS improve genome coverage and integrate DNA methylation with gene expression analysis in single cells. The technique revealed new insights into methylation patterns in maize microspores, suggesting its use in any species18.
Long-read bisulfite sequencing technology improvements
Long-read sequencing technologies offer advantages in detecting base modifications and provide longer reads for better mapping.
Double-strand bisulfite sequencing (DSBS) enables the simultaneous detection of single-nucleotide variants and DNA methylation, offering a cost-effective approach for genetic and epigenetic studies. Given that the approach could detect the hemimethylation patterns in human cells, DSBS can be applied to augment our understanding of diseases19.
Integration with multi-omics data for comprehensive epigenetic studies
Integrating bisulfite sequencing data with other omics layers, such as transcriptomics and proteomics, helps to understand the complex interactions that are involved in the development of disease.
For example, whole genome bisulfite sequencing and transcriptome sequencing in high-grade serous ovarian cancer samples helped identify that expression patterns were conserved in primary tumors despite chemotherapy, providing insights into chemoresistance20.
Techniques like smart-reduced representation bisulfite sequencing (Smart-RRBS) demonstrate the potential of multi-omics integration for uncovering insights into epigenetic modifications. This method combines Smart-seq2 (a single-cell RNA sequencing protocol) with RRBS (a targeted bisulfite sequencing approach) to enable simultaneous measurement of DNA methylation and gene expression in the same cell.
Smart-RRBS bridges the gap between bulk RRBS (cost-effective but lacking single-cell resolution) and WGBS (comprehensive but expensive), making it a powerful tool for exploring epigenetic and transcriptional interplay in heterogeneous cell populations21.
Future potential of bisulfite sequencing in research and personalized medicine
Bisulfite sequencing has the potential for personalized medicine, with advancements like scBS-seq offering precise insights into drug discovery. The continuous refinement of bisulfite sequencing protocols, along with the integration of long-read and multi-omics technologies, will enhance our understanding of epigenetic regulation and its implications for disease treatment and prevention22.
Recent advancements in sequencing technologies have enabled high-quality DNA sequencing for various applications, including non-invasive disease detection, cancer screening, and profiling DNA methylation patterns in both clinical and research settings. Bisulfite sequencing consistently delivers reliable results in clinical sample analysis, underscoring its potential for early cancer detection through cfDNA methylation profiling.
Bisulfite sequencing remains a cornerstone technique for studying DNA methylation. Despite its challenges, ongoing advancements promise to enhance its utility in research and medicine, paving the way for better understanding and treatment of various diseases.
FAQs
What are the main steps involved in bisulfite sequencing?
Bisulfite sequencing involves three main steps:
1. Treating DNA with sodium bisulfite to convert unmethylated cytosines to uracils while methylated cytosines remain unchanged.
2. Amplifying the treated DNA by PCR to generate sufficient quantities for sequencing.
Sequencing the amplified DNA and comparing the treated DNA sequence to the original sequence to identify methylated cytosines (those that remained as cytosines) and unmethylated cytosines (those converted to thymines).
How does bisulfite sequencing differ from other DNA methylation analysis techniques?
Bisulfite sequencing provides single-base resolution of DNA methylation by directly sequencing treated DNA, offering detailed analysis of methylation patterns. In contrast, other techniques like methylation-specific PCR and microarrays are less precise, focusing on specific regions or a broader overview of methylation without base-specific detail.
What are the advantages of using high-throughput sequencing in bisulfite sequencing?
High-throughput sequencing in bisulfite sequencing allows for the simultaneous analysis of methylation across the entire genome, providing a comprehensive view of DNA methylation patterns. It increases the speed and scale of data collection, offers single-base resolution, and provides accurate quantification of methylation levels.
How does single-cell bisulfite sequencing improve our understanding of DNA methylation?
Single-cell bisulfite sequencing allows the analysis of DNA methylation at the individual cell level, revealing heterogeneity in methylation patterns across different cells. This technique provides insights into cell-specific methylation changes, helping to understand cellular diversity and the role of methylation in development and disease.
What are the limitations of pyrosequencing in bisulfite sequencing?
Pyrosequencing in bisulfite sequencing has limited throughput and is typically restricted to analyzing shorter DNA regions, which can reduce the scope of methylation analysis. It can be sensitive to sequencing errors and may not accurately detect methylation in regions with high GC content. Additionally, the high number of PCR cycles required can result in a high standard deviation in methylation quantification.
References
1. Ludgate, J.L., Wright, J., Stockwell, P.A., et al. A streamlined method for analysing genome-wide DNA methylation patterns from low amounts of FFPE DNA. BMC Med Genomics. 10, 54 (2017).
2. Zhang, S., He, S., Zhu, X., et al. DNA methylation profiling to determine the primary sites of metastatic cancers using formalin-fixed paraffin-embedded tissues. Nat Commun. 14, 5686 (2023).
3. Lam, D., Luu, P.L., Song, J.Z., et al. Comprehensive evaluation of targeted multiplex bisulphite PCR sequencing for validation of DNA methylation biomarker panels. Clin Epigenet. 12, 90 (2020).
4. Gong, T., Borgard, H., Zhang, Z., et al. Analysis and performance assessment of the whole genome bisulfite sequencing data workflow: currently available tools and a practical guide to advance DNA methylation studies. Small Methods. 6, e2101251 (2022).
5. Nakabayashi, K., Yamamura, M., Haseagawa, K., et al. Reduced representation bisulfite sequencing (RRBS). Methods Mol Biol. 2577, 39–51 (2023).
6. Moser, D.A., Müller, S., Hummel, E.M., et al. Targeted bisulfite sequencing: a novel tool for the assessment of DNA methylation with high sensitivity and increased coverage. Psychoneuroendocrinology. 120, 104784 (2020).
7. De Borre, M., Branco, M.R., et al. Oxidative bisulfite sequencing: an experimental and computational protocol. Methods Mol Biol. 2198, 333–348 (2021).
8. Leontiou, C.A., Besa, E., Dimou, N.L., et al. Bisulfite conversion of DNA: performance comparison of different kits and methylation quantitation of epigenetic biomarkers that have the potential to be used in non-invasive prenatal testing. PLoS One. 10, e0135058 (2015).
9. Foox, J., Nordlund, J., Lalancette, C., et al. The SEQC2 epigenomics quality control (EpiQC) study. Genome Biol. 22, 332 (2021).
10. Raine, A., Yang, W., Thompson, K., et al. Data quality of whole genome bisulfite sequencing on Illumina platforms. PLoS One. 13, e0195972 (2018).
11. Welsh, H., Batalha, C.M.P.F., Li, W., et al. A systematic evaluation of normalization methods and probe replicability using Infinium EPIC methylation data. Clin Epigenet. 15, 41 (2023).
12. Tost, J., Ak-Aksoy, S., Campa, D., et al. Leveraging epigenetic alterations in pancreatic ductal adenocarcinoma for clinical applications. Semin Cancer Biol. 109, 101–124 (2025).
13. Legendre, C., Gooden, G.C., Johnson, K., et al. Whole-genome bisulfite sequencing of cell-free DNA identifies signature associated with metastatic breast cancer. Clin Epigenet. 7, 100 (2015).
14. Dahlet, T., Argüeso Lleida, A., Al Adhami, H., et al. Genome-wide analysis in the mouse embryo reveals the importance of DNA methylation for transcription integrity. Nat Commun. 11, 3153 (2020).
15. Duan, J.E., Jiang, Z.C., Alqahtani, F., et al. Methylome dynamics of bovine gametes and in vivo early embryos. Front Genet. 10, 512 (2019).
16. Chatterton, Z., Mendelev, N., Chen, S., et al. Bisulfite amplicon sequencing can detect glia and neuron cell-free DNA in blood plasma. Front Mol Neurosci. 14, 672614 (2021).
17. Zhang, Q., Ma, S., Liu, Z., et al. Droplet-based bisulfite sequencing for high-throughput profiling of single-cell DNA methylomes. Nat Commun. 14, 4672 (2023).
18. Li, X., Chen, L., Zhang, Q., et al. BRIF-Seq: bisulfite-converted randomly integrated fragments sequencing at the single-cell level. Mol Plant. 12, 438–446 (2019).
19. Liang, J., Zhang, K., Yang, J., et al. A new approach to decode DNA methylome and genomic variants simultaneously from double strand bisulfite sequencing. Brief Bioinform. 22, bbab201 (2021).
20. Gull, N., Fan, H., Singh, S., et al. DNA methylation and transcriptomic features are preserved throughout disease recurrence and chemoresistance in high grade serous ovarian cancers. J Exp Clin Cancer Res. 41, 232 (2022).
21. Gu, H., Raman, A.T., Wang, X., et al. Smart-RRBS for single-cell methylome and transcriptome analysis. Nat Protoc. 16, 4004–4030 (2021).
22. Choi, Y., Choi, D.W., Lee, S. Multi-omics techniques for the genetic and epigenetic analysis of rare diseases. J Genet Med. 20, 1–5 (2023).