What Is Whole Genome Sequencing and How Does It Work?

Published: January 24, 2026 | Author: Editorial Team | Last Updated: January 24, 2026
Published on humansgenomes.com | January 24, 2026

Whole genome sequencing (WGS) is the process of determining the complete DNA sequence of an organism's genome — every one of the approximately 3 billion base pairs in the human genome — in a single experiment. Once an extraordinarily expensive and slow endeavour, WGS has been transformed by advances in sequencing technology into a routine clinical and research tool whose cost continues to fall. Understanding how it works and what it can reveal is increasingly important for clinicians, researchers, and anyone interested in the future of personalised medicine.

From DNA to Sequenceable Fragments

The sequencing process begins with extracting DNA from a biological sample — typically blood, saliva, or tissue. The extracted DNA is then fragmented into shorter pieces, typically a few hundred to a few thousand base pairs long, depending on the sequencing platform. These fragments are processed in a library preparation step that attaches sequencing adapters — short synthetic sequences — to the ends of each fragment. These adapters allow the fragments to bind to the sequencing instrument and be amplified before reading.

Sequencing by Synthesis: The Dominant Technology

The most widely used WGS technology today is Illumina sequencing, based on a process called sequencing by synthesis (SBS). Prepared DNA fragments are loaded onto a flow cell — a glass chip whose surface is studded with molecules that capture individual fragments and amplify them into small clusters. The sequencer then reads each cluster by adding fluorescently labelled nucleotides one at a time; each nucleotide emits a specific colour as it is incorporated, and a camera records the sequence of colour signals, translating them into a base sequence. A typical WGS run produces billions of short "reads" of 100 to 300 base pairs each.

Long-Read Sequencing Technologies

Illumina's short-read approach dominates the market for accuracy and throughput, but long-read platforms from Pacific Biosciences (PacBio) and Oxford Nanopore Technologies are increasingly important for regions of the genome that short reads cannot resolve. Long reads — spanning tens of thousands to hundreds of thousands of base pairs — can traverse highly repetitive genomic regions, phasecomplex structural variants, and resolve the full-length structure of alternative splice isoforms. The T2T-CHM13 complete human genome sequence — the first truly gapless human genome — was produced using long-read technology, demonstrating its capacity to access regions that short reads cannot reliably assemble.

From Raw Reads to Biological Insight

Raw sequencing output must be processed through a computational pipeline before it becomes biologically interpretable. Reads are aligned to a reference genome, and computational tools identify positions where the sample's sequence differs from the reference — calling variants including SNPs, insertions, deletions, and structural variants. These variants are then annotated against databases of known variants, predicted for functional impact, and interpreted in the clinical or research context of interest. For a clinical WGS report, this process typically concludes with a clinical geneticist reviewing computationally flagged variants and writing an interpretive report.

Explore our whole genome sequencing resources and population genomics tools, or contact us for information about current sequencing technologies and applications.

← Back to Home

Subscribe to Our Newsletter

Join 10,000+ subscribers. Get the latest updates, exclusive content, and expert insights delivered to your inbox weekly.

No spam. Unsubscribe anytime. We respect your privacy.