Sanger sequencing: Principles, process and applications
Sanger sequencing, also called the “chain termination method,” is a first-generation sequencing technique widely used in molecular biology for DNA sequencing.
The order of nucleic acids in DNA is important for the hereditary and biochemical functions and properties of life. Hence, DNA sequencing becomes vital to understand these crucial aspects; first-generation DNA sequencing includes Maxam-Gilbert (chemical degradation) and Sanger (chain termination); second-generation sequencing facilitates the quick sequencing of whole genomes; and third-generation sequencing methods sequence in real-time without amplification.
Sanger sequencing involves the selective incorporation of chain-terminating dideoxynucleotides (ddNTPs) during DNA replication with the help of a DNA polymerase enzyme. It is highly accurate and reliable, making it ideal for small DNA sequences or regions requiring precision. Its applications include mutation detection, plasmid verification, and Polymerase chain reaction (PCR) product analysis. It remains an essential tool for validating results from high-throughput sequencing methods.
Historical background of Sanger sequencing
Frederick Sanger, often called the “father of genomics,” was a two-time Nobel laureate in chemistry. His groundbreaking work on the “dideoxy” chain-termination method led to the first widely applicable DNA sequencing technique, earning him his second Nobel prize. This method became the foundation of Sanger sequencing, which was developed in the 1970s and first used in 1977 to sequence the genome of a bacteriophage.
Sanger sequencing revolutionized genetic research and remained the gold standard for over three decades. Its impact extended to major scientific initiatives, including the Human Genome Project, where the Wellcome Sanger Institute, founded in 1993, played a pivotal role.
Sequencing technologies have evolved rapidly, beginning with manual Sanger sequencing and advancing to automated methods in the 1980s. Key milestones in this evolution include the introduction of fluorescent labels, the development of pyrosequencing, and the advent of next-generation sequencing (NGS) platforms.
These advancements significantly reduced costs and increased sequencing speeds, making ambitious projects like the Human Genome Project and the 100,000 Genomes Project possible. Today, modern sequencing technologies continue to push the boundaries of biomedical research, with applications spanning from personalized medicine to disease detection.
Principles of Sanger sequencing
Sanger sequencing works on the principle of chain termination during the synthesis of new DNA strands. This is achieved by incorporating ddNTPs into the reaction, which prevents further elongation. By including fluorescently labeled ddNTPs in a DNA polymerase reaction, DNA fragments of varying lengths are generated, each corresponding to the sequence of the DNA template. The resulting fragments are then separated by gel electrophoresis to determine the DNA sequence.
This image is from an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms. Reference: Nafea, A. M., Wang, Y., Wang, D., Salama, A. M., Aziz, M. A., Xu, S., & Tong, Y. (2024). Application of next-generation sequencing to identify different pathogens. Frontiers in microbiology, 14, 1329330. https://doi.org/10.3389/fmicb.2023.1329330
Figure: The Sanger sequencing method includes the following steps: denaturing dsDNA, making multiple copies, primer attachment, polymerase solution addition, chain amplification, denaturation, and electrophoreses. The extension of the DNA strand is terminated when fluorescently labeled ddNTP is incorporated into the DNA strand; electrophoresis is used to separate the fragments, and the fluorescence intensity of different colors after laser irradiation is used to derive the sequence.
Steps in the Sanger sequencing process
The Sanger sequencing process involves a series of precise steps to determine the nucleotide sequence of DNA, starting with template preparation and culminating in data analysis.
DNA template preparation
The target DNA is extracted and purified to prepare a single-stranded DNA template. Methods like chemical or column-based extractions are commonly used for isolating DNA1.
DNA polymerase reaction
PCR amplifies the target DNA, incorporating normal dNTPs and small quantities of ddNTPs. Each ddNTP is uniquely labeled with fluorescent markers to identify the terminating base.
Chain termination and fragment generation
The ddNTPs randomly terminate the elongation process, producing DNA fragments of different lengths. Each fragment ends with a ddNTP corresponding to one of the four nucleotide bases.
Gel electrophoresis and capillary electrophoresis
DNA fragments are then separated based on size using gel or capillary electrophoresis. In modern methods, capillary electrophoresis offers high-resolution separation in a single reaction tube, making the process more efficient and precise.
Data analysis and sequence assembly
Fluorescent signals from the terminated fragments are detected to identify the nucleotide at each position. The sequence is determined from the order of fluorescence peaks in a chromatogram and then assembled for final analysis.
This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Reference: Shaibu JO, Onwuamah CK, James AB, Okwuraiwe AP, Amoo OS, Salu OB, et al. (2021) Full length genomic sanger sequencing and phylogenetic analysis of Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) in Nigeria. PLoS ONE 16(1): e0243271. https://doi.org/10.1371/journal.pone.0243271
Figure: A sample sequencing chromatogram: snapshots of N-region and ORF1ab region of the assembled individual sequence chromatograms from a study on full-length genomic sanger sequencing and phylogenetic analysis of Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) in Nigeria.
Sanger sequencing process
Sanger sequencing relies on four key components, each playing a specific role in ensuring accurate and efficient determination of DNA sequences. The process typically begins with samples such as plasmid DNA, cell lines, and lysates, which are prepared to isolate and amplify the target DNA for sequencing.
Role of ddNTPs
ddNTPs are modified nucleotides that lack the 3'-hydroxyl group required for forming a phosphodiester bond. Their incorporation into a growing DNA strand terminates the chain, enabling sequence determination by generating fragments of varying lengths.
Chain-terminating mechanism
The chain-termination mechanism relies on ddNTPs being randomly incorporated during DNA synthesis. Once a ddNTP is added, the lack of a 3'-OH prevents further elongation, effectively halting DNA synthesis at specific nucleotides.
DNA polymerase
DNA polymerase catalyzes the synthesis of a complementary DNA strand by adding nucleotides to a primer. It incorporates both standard dNTPs for elongation and ddNTPs for random termination, creating a mixture of terminated DNA fragments.
Primers in Sanger sequencing
Primers are short, single-stranded DNA sequences that bind specifically to the template DNA. They provide a starting point for DNA polymerase, ensuring targeted and directional synthesis during the sequencing process.
Sanger sequencing vs. NGS
When comparing Sanger sequencing and NGS, key differences emerge in terms of throughput, accuracy, cost, and read length, each influencing their suitability for various genomic applications.
Sanger sequencing is ideal for analyzing small DNA regions or a limited number of genomic targets, typically fewer than 20, due to its high accuracy and straightforward workflow. In contrast, NGS is more effective for large-scale projects. It allows the simultaneous and cost-effective analysis of hundreds to thousands of genes, making it suitable for studies requiring high throughput, sensitivity, and discovery power.
Sanger sequencing vs. PCR
PCR is a versatile method used to replicate specific DNA regions, producing millions of copies from a minimal starting amount of material. The process relies on the cyclical process of denaturation, annealing, and extension, catalyzed by a heat-stable DNA polymerase.
Differences in purpose and methodology
Sanger sequencing is ideal for determining the precise sequence of a DNA fragment, especially when identifying mutations or confirming genetic constructs. PCR is best used when the goal is to amplify a specific DNA target for further applications, such as cloning, analysis, or preparation for sequencing.
Applications of Sanger sequencing
The Sanger sequencing method is renowned for its accuracy and reliability. It is a versatile tool widely applied across diverse fields of research and clinical practice.
Clinical diagnostics and genetic testing
Sanger sequencing is widely used in both research and clinical settings for its high accuracy in detecting single nucleotide variants and small insertions/deletions. It is commonly employed for diagnostic sequencing of single genes and identifying specific familial sequence variants, such as those linked to conditions like BRCA1-related breast cancer or autosomal recessive disorders like cystic fibrosis. This technique is also essential for prenatal testing, carrier screening, and segregation analysis to evaluate the pathogenicity of variants.
Microbial identification and infectious disease studies
Sanger sequencing plays a pivotal role in microbial identification and infectious disease studies by enabling precise analysis of genetic sequences. It is commonly used for analyzing 16S rRNA genes to identify bacterial genera and species, providing detailed insights into microbial phylogeny. This method is particularly effective in identifying single clones from pure cultures and verifying species-level identification in clinical samples.
Validation of NGS results
Sanger sequencing is vital for validating NGS results and ensuring the accuracy and reliability of variant identification. It is often used to confirm clinically significant variants, especially in complex genomic regions such as AT-rich, GC-rich sequences, or pseudogenes, where NGS may produce false positives. By amplifying target regions and sequencing them with high precision, Sanger sequencing is a complementary and standard reference to resolve discrepancies and refine NGS data.
Other applications
Sanger sequencing is essential for confirming the sequence integrity of mRNAs used in vaccine and therapeutic manufacturing, ensuring they meet the stringent quality and safety standards of regulatory bodies. Its high accuracy and reproducibility are vital in detecting any unanticipated sequence variations during plasmid production and mRNA identity testing.
With a streamlined workflow that eliminates the need for PCR amplification, Sanger sequencing provides rapid and unambiguous results, enabling efficient quality control. Sanger sequencing is widely used for mitochondrial DNA analysis due to its high accuracy and ability to generate long-read sequences. This makes it especially effective for studying mutations and variations linked to human diseases.
It helps researchers explore the role of mitochondrial DNA in aging, diabetes, certain cancers, and other conditions by providing precise insights into genetic changes. Its reliability also extends to applications in population genetics, forensics, and biodiversity research.
Advantages and limitations of Sanger sequencing
Key advantages:
- Excellent for detecting single nucleotide variants and small insertions/deletions.
- Provides precise, base-by-base sequencing, ensuring highly reliable results.
- Generates reads up to 900 bp, reducing the need for overlapping sequences.
- Ideal for sequencing complex or repetitive regions.
- Useful for validating NGS results and targeted gene sequencing.
- Eliminates extensive computational assembly in small-scale studies, making it efficient and reliable.
Key limitations:
- Processes single DNA fragments at a time, making it unsuitable for large-scale genome sequencing.
- Requires manual preparation, leading to slower output compared to high-throughput methods like NGS.
- Expensive and labor-intensive for whole-genome sequencing.
- Best suited for small-scale sequencing tasks or validating results from high-throughput platforms.
Trends in Sanger sequencing
The advancement and integration of DNA sequencing technologies have improved genetic research, enabling precise analyses, validation, and discovery of genetic variants.
Integrating with NGS
Integrating Sanger sequencing with NGS combines the high accuracy of Sanger sequencing with the efficiency and scalability of NGS. This integration is particularly valuable for identifying and validating genetic variants in clinical and research applications.
For example, a study showed that Sanger sequencing data could be easily integrated with NGS-derived whole-genome sequencing data where targeted Sanger sequencing was a quick and cost-effective approach to increase the whole genome coverage2. This becomes significant when monitoring new variants of pathogens like severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), where genomic regions cannot be “missed.”
Advancements and potential improvements
First-generation sequencing, exemplified by Sanger sequencing, remains a gold standard for its high accuracy and long-read lengths. While newer sequencing technologies have introduced higher throughput and lower costs, Sanger sequencing is still widely used for targeted sequencing applications, such as validating NGS results, sequencing plasmids post-cloning, analyzing mitochondrial DNA, and confirming clinically significant variants.
Recent trends in Sanger sequencing focus on improving automation and integration with modern technologies, making the process faster and more cost-effective. For instance, advancements in capillary electrophoresis and fluorescent labeling techniques have enhanced resolution and accuracy.
Despite being considered a first-generation method, Sanger sequencing continues to play an important role in precision applications, such as mutation detection, plasmid verification, and high-resolution analysis of complex genomic regions.
Role in genetic research
Sanger sequencing is essential for detecting single nucleotide polymorphisms and small indels with high accuracy, providing a robust foundation for genetic studies. It enables the precise validation of genetic variants identified by NGS, ensuring data reliability.
It also excels in sequencing challenging regions, such as those with repetitive elements, by generating clear chromatograms. Advances in troubleshooting technical artifacts further refine their application in clinical and molecular research.
FAQs
What is Sanger sequencing?
Sanger sequencing is a DNA sequencing method developed in 1970 that uses chain-terminating ddNTPs to create DNA fragments of varying lengths during replication. These fragments are separated by electrophoresis, and their sequence is determined by detecting fluorescent or radiolabeled ddNTPs. It is a highly accurate method still used for small-scale studies and validation purposes.
What are the advantages of using Sanger sequencing for repetitive regions of the genome?
Sanger sequencing is highly effective for analyzing repetitive regions of the genome due to its ability to produce long, high-quality reads of up to 900 bp, reducing the ambiguity associated with short-read sequencing methods (paired-end read lengths of 90- 150bp). This precision enables researchers to resolve complex repeats and accurately identify variations in challenging genomic areas.
How has Sanger sequencing evolved over the years?
Sanger sequencing has evolved significantly since its inception, transitioning from labor-intensive manual processes using radiolabeled ddNTPs and polyacrylamide gels to more efficient and automated systems with fluorescent labeling and capillary electrophoresis.
These advancements improved accuracy, read lengths, and throughput, making the technique a reliable standard for DNA analysis while adapting to complement high-throughput next-generation sequencing technologies.
References
-
Crossley, B.M., Bai, J., Glaser, A., et al. Guidelines for Sanger sequencing and molecular assay monitoring. J Vet Diagn Invest. 32, 767-775 (2020).
Singh, L., San, J.E., Tegally, H., et al. Targeted Sanger sequencing to recover key mutations in SARS-CoV-2 variant genome assemblies produced by next-generation sequencing. Microb Genom. 8, (2022).