
Whole-genome sequencing, or WGS, reads DNA across coding and noncoding regions instead of capturing mainly exons. A clinical genome can be analyzed for single-base changes, small insertions and deletions, copy-number variants, structural rearrangements, mitochondrial variants, and selected repeat expansions. This broad scope can reveal diagnoses missed by gene panels or exome sequencing, particularly when the causal change lies deep within an intron, disrupts chromosome structure, or is difficult to detect with capture-based methods.
WGS is not a perfect scan of every genetic mechanism. Most current clinical tests use short reads, which remain challenging in highly repetitive DNA, segmental duplications, centromeres, and some large repeat expansions. Laboratories also differ in which variant types and genomic regions they analyze and report. A genome may be sequenced broadly but interpreted through a limited virtual gene panel. The most useful result comes from matching the laboratory’s validated pipeline to the clinical question and combining genomic findings with phenotype, family history, and functional evidence.
- WGS samples nearly the entire nuclear genome and often mitochondrial DNA in one sequencing workflow.
- It can detect coding, intronic, copy-number, and structural variants when the laboratory validates those analyses.
- Trio sequencing with parents can improve interpretation of rare-disease cases.
- A negative genome does not exclude every repeat, mosaic, epigenetic, or technically difficult variant.
- Stored genome data can be reanalyzed as new disease genes and analytic methods become available.
Table of Contents
- What Whole-Genome Sequencing Covers
- Short-Read and Long-Read Genomes
- The Clinical WGS Workflow
- Variant Types a Genome May Detect
- Interpreting Positive, Uncertain, and Negative Results
- Blind Spots, Data Scope, and Secondary Findings
- Confirmation, Reanalysis, and Next Steps
What Whole-Genome Sequencing Covers
The human genome contains about three billion DNA bases. Protein-coding exons account for only a small fraction. The remaining sequence includes introns, promoters, enhancers, untranslated regions, noncoding RNA genes, repetitive elements, centromeric and telomeric DNA, and regions whose functions are still being defined.
WGS avoids the exon-capture step used in whole-exome sequencing. DNA fragments are prepared and sequenced directly across the genome. This generally produces more even coverage than exome capture and allows one dataset to support analysis of many variant classes.
Clinical WGS may be useful when:
- a rare inherited disorder is suspected but many genes could be involved;
- exome or panel testing was negative or found only one variant in a recessive gene;
- a structural rearrangement, repeat, or deep intronic variant is possible;
- several test types would otherwise be ordered sequentially;
- a critically ill infant needs rapid broad testing;
- a complex phenotype suggests more than one diagnosis;
- reanalysis over time is likely to be valuable.
“Whole genome” describes where sequence data are generated, not necessarily what is interpreted. Some laboratories use virtual panels and initially review only genes associated with the patient’s phenotype. Others analyze the entire known disease genome. Noncoding variants may be reported only in genes with established evidence and validated interpretation rules.
This distinction protects against overwhelming numbers of uncertain findings. Every genome contains millions of variants, and most noncoding changes cannot yet be interpreted clinically. Broad data generation therefore coexists with focused analysis.
Genome sequencing can reduce the need for separate tests, but it does not automatically replace karyotype, microarray, repeat-expansion assays, methylation studies, RNA analysis, or biochemical testing. Whether it functions as a “one-stop” test depends on the laboratory’s validated variant callers, coverage, reporting policy, and the suspected disorder.
Diagnostic yield varies. A person with a strongly monogenic phenotype and informative family samples has a different chance of diagnosis from someone with nonspecific symptoms and extensive prior testing. Studies also differ in whether WGS is first line, follows a negative exome, or includes research-level analysis. A published yield should not be presented as a guaranteed personal outcome.
Short-Read and Long-Read Genomes
Most routine clinical WGS currently uses short-read sequencing. Long-read technologies are increasingly available and can solve variants that short reads cannot, but they have different requirements and validation considerations.
Short-read WGS
Short-read platforms produce hundreds of millions of paired DNA fragments, commonly around 100 to 250 bases each. Software aligns these reads to a reference genome. Typical germline testing aims for an average depth that allows most positions to be read many times.
Short reads provide high base accuracy and efficient detection of single-nucleotide variants and small indels in unique sequence. They also support copy-number and structural variant analysis through read depth, unexpectedly spaced pairs, split reads, and changes in allele balance.
Their main weakness is ambiguity. A short sequence may match several locations in a repetitive gene, pseudogene, segmental duplication, or mobile element. Reads can be misaligned or discarded, leaving gaps and making breakpoints difficult to reconstruct.
Long-read WGS
Long-read systems produce reads thousands to hundreds of thousands of bases long. A read can span an entire repeat, gene conversion, structural breakpoint, or region containing several variants. This helps with phasing—determining which variants are on the same chromosome—and with building more complete personal genome assemblies.
Long-read WGS may improve detection of:
- large insertions, inversions, and complex structural variants;
- tandem-repeat expansions and repeat interruptions;
- variants in genes with pseudogenes or segmental duplications;
- mobile element insertions;
- methylation patterns on platforms that measure base modifications;
- haplotypes and compound heterozygous phase.
Clinical adoption requires validated accuracy, standardized quality measures, robust software, high-quality long DNA molecules, and interpretable reference databases. Long-read sequencing is not automatically comprehensive. Coverage can be uneven, analysis is computationally demanding, and evidence for some newly visible variants remains limited.
Genome references and pangenomes
Traditional analysis aligns reads to one linear reference genome. That reference does not represent all human structural diversity equally. Pangenome references incorporate multiple haplotypes and may improve alignment and variant discovery, particularly across diverse ancestries and complex regions. Clinical use is developing and requires stable nomenclature and validated interpretation.
The question “short read or long read?” should be answered by the suspected variant. Short-read WGS remains highly effective for many rare diseases. Long reads are especially valuable after an unresolved result points to repeat, phase, pseudogene, methylation, or structural complexity.
The Clinical WGS Workflow
WGS begins before the sequencer. The ordering team defines the phenotype, collects family history, discusses possible results, and decides whether to include parents or other relatives.
Consent and test design
Consent commonly addresses the primary diagnostic question, secondary findings, unexpected family relationships, data storage, reanalysis, and whether de-identified data can be used for research. Policies differ for children, prenatal testing, and critically ill patients.
A singleton tests only the patient. A duo adds one relative. A trio includes both biological parents and is often preferred for severe early-onset disease. Trio data can identify de novo variants, establish phase in recessive disorders, and filter inherited benign variation.
Specimen and DNA extraction
Blood is commonly used because it yields high-quality DNA. Saliva or cheek swabs may be accepted. Long-read sequencing may require ultra-high-molecular-weight DNA and special handling. Another tissue should be considered when blood may not contain the variant, as in tissue-limited mosaic disorders.
People who have received an allogeneic bone marrow transplant may have donor-derived blood DNA. A skin fibroblast or other nonhematologic specimen may be required for germline testing. Tumor WGS uses separate workflows and matched normal DNA when possible.
Library preparation and sequencing
DNA is fragmented for short-read WGS or preserved in long molecules for long-read testing. Adapters and sample identifiers are added. The instrument generates raw sequence signals, which are converted to base calls and quality scores.
Coverage is described by average depth, uniformity, and the percentage of the genome meeting a minimum threshold. A high average does not ensure every medically relevant base is adequately covered. Laboratories monitor contamination, sample sex consistency, relatedness, and other quality metrics.
Bioinformatic analysis
Reads are aligned to a reference, duplicates and technical artifacts are addressed, and variant callers generate candidate files. Different algorithms search for different event types. A comprehensive workflow may include:
- single-nucleotide and small-indel calling;
- copy-number analysis;
- structural variant calling;
- mitochondrial variant analysis;
- repeat-expansion screening;
- regions of homozygosity and uniparental disomy analysis;
- mobile element detection;
- mosaic variant calling;
- pharmacogenetic or HLA analysis when validated.
Not every clinical genome includes each item. The report should state what was analyzed.
Prioritization and interpretation
Millions of differences are filtered using frequency, predicted consequence, inheritance, gene-disease validity, and phenotype match. Analysts may begin with a virtual gene panel and expand to the broader genome. Candidate noncoding variants require strong evidence that the region regulates a relevant gene or affects splicing.
Multidisciplinary review can be valuable for complex cases. Clinical geneticists, laboratory scientists, bioinformaticians, radiologists, pathologists, and disease specialists may combine information that no single data source resolves.
Reporting and turnaround
Routine WGS often takes several weeks. Rapid genome sequencing can produce preliminary or final results in days for critically ill newborns and children. Rapid testing prioritizes variants with immediate management implications and may issue updated reports as analysis continues.
Variant Types a Genome May Detect
The central advantage of WGS is the ability to examine several variant classes from one dataset. Actual performance depends on read length, coverage, software, and validation.
Single-nucleotide variants and small indels
Short-read WGS is highly reliable for single-base changes and many small insertions and deletions in unique regions. It covers coding exons without capture gaps and also identifies intronic and regulatory candidates. Difficult homopolymers and repetitive regions remain challenging.
Deep intronic and regulatory variants
A deep intronic variant can create a new splice site or pseudoexon, causing abnormal RNA to include sequence that should have been removed. WGS can locate the DNA change, but proving its effect often requires RNA testing, minigene assays, or other functional evidence.
Promoter and enhancer variants are harder to interpret because regulatory elements can act over long distances and differ by tissue. Laboratories generally report only those with established gene-specific evidence or compelling functional support.
Copy-number variants
Read depth can identify deletions and duplications from part of an exon to large chromosome regions. Genome data can often define breakpoints more precisely than microarray and may reveal where duplicated material is inserted. Low-complexity regions and mosaic events can reduce sensitivity.
Structural variants
Paired and split reads can detect inversions, translocations, insertions, mobile elements, and complex rearrangements. WGS can expose a gene disrupted by a breakpoint even when no DNA is gained or lost. Short reads may not fully resolve events embedded in repeats, while long reads can span them more directly.
A clinical structural variant test may still be needed for confirmation, chromosome context, or mosaic cell counting.
Tandem-repeat expansions
Specialized software can screen short-read WGS data for selected repeat expansions. Some alleles can be sized accurately; others are only flagged as expanded because the repeat is longer than the read. Very large, GC-rich, or complex repeats may require repeat-primed PCR, Southern blot, optical mapping, or long-read sequencing.
A negative genome report does not exclude repeat disease unless the specific locus and size range were validated.
Mitochondrial DNA
WGS typically generates substantial mitochondrial coverage. A validated pipeline can detect mitochondrial single-nucleotide variants and some deletions. Heteroplasmy thresholds and tissue differences are critical. Blood levels can decline with age for some variants and may not represent muscle or urine.
Regions of homozygosity and uniparental disomy
Genome-wide genotypes can identify long homozygous segments that suggest parental relatedness or uniparental disomy. Sequence data may also show a recessive variant within such a region. Methylation testing may be needed to confirm an imprinting disorder.
Mosaic variants
High-quality WGS can detect some variants present in a fraction of cells, but standard depth may be inadequate for low-level mosaicism. Deep targeted sequencing of affected tissue is often more sensitive. Mosaic structural variants may require cytogenetics, microarray, or optical mapping.
Phasing
Parental data and long reads help determine whether variants occur in cis or trans. Phase is crucial in recessive disease and can clarify complex pharmacogenetic alleles. Short-read statistical phasing is less reliable for rare variants far apart.
Interpreting Positive, Uncertain, and Negative Results
A genome report is organized around clinical relevance rather than listing every detected variant.
Diagnostic result
A diagnostic result identifies pathogenic or likely pathogenic variant(s) that fit the phenotype and inheritance. The finding may be a coding variant, structural change, repeat expansion, mitochondrial variant, or a noncoding change with convincing functional evidence.
One genome can reveal dual diagnoses. For example, one variant may explain a neurologic condition while another explains an unrelated renal finding. The report should distinguish which features are accounted for and which remain unexplained.
A positive result can guide treatment, surveillance, prognosis, reproductive counseling, and family testing. The clinical effect is gene specific. Not all molecular diagnoses have a targeted therapy, and some provide mainly prognostic or reproductive information.
Partial or candidate result
The laboratory may find one pathogenic variant in a recessive gene but not the second, or a strong candidate noncoding variant without sufficient functional evidence. This is not a complete diagnosis. Follow-up may include RNA sequencing, long-read analysis, deletion testing, or testing another tissue.
Candidate gene findings may be shared through research matching services. A rare damaging variant in a gene with no established human disease relationship should not be presented as clinically confirmed without additional evidence.
Variant of uncertain significance
A VUS has insufficient or conflicting evidence. Broad genome analysis can generate many uncertain candidates, particularly in noncoding regions. Clinical laboratories limit reporting to avoid overwhelming patients with variants that cannot guide care.
A VUS should not be used alone for preventive surgery, pregnancy decisions, or predictive testing of healthy relatives. Segregation, functional studies, and reanalysis may later change the classification.
Negative result
A negative result means no reportable explanation was identified within the analyzed scope. It does not mean the genome is free of variants or that the condition is not genetic.
Possible explanations include:
- the causal variant is in a region with poor mapping or coverage;
- the repeat or structural event exceeds the pipeline’s capability;
- mosaicism is absent from blood or below the detection threshold;
- the mechanism is epigenetic, transcriptomic, or biochemical rather than detectable DNA sequence;
- the responsible gene has not yet been linked to disease;
- the variant was filtered because the phenotype information was incomplete;
- the condition is multifactorial or nongenetic.
A negative WGS can still reduce the likelihood of some disorders and prevent redundant sequencing. Its value depends on the validated analyses and the pretest clinical hypothesis.
Blind Spots, Data Scope, and Secondary Findings
The term whole genome can create unrealistic expectations. Some portions of the genome remain difficult to read or interpret.
Short-read blind spots include highly repetitive sequence, segmental duplications, paralogous genes, centromeres, telomeres, acrocentric short arms, and large tandem repeats. Newer references and long reads improve access, but clinical validation and variant databases lag behind technical discovery.
Variant classes that may require separate testing include:
- very large or complex repeat expansions;
- balanced rearrangements with breakpoints in repetitive DNA;
- low-level mitochondrial heteroplasmy;
- tissue-specific somatic mosaicism;
- methylation and imprinting abnormalities;
- epigenetic signatures;
- RNA expression and splicing effects;
- some gene conversions and pseudogene-associated events.
Analytic scope matters as much as sequencing. A laboratory may sequence the entire genome but report only variants in a phenotype panel. Another may perform open genome-wide analysis. The test description should identify:
- genes or regions reviewed;
- minimum coverage and excluded regions;
- structural and copy-number size limits;
- repeat loci screened;
- mitochondrial analysis and heteroplasmy threshold;
- mosaic detection threshold;
- noncoding reporting policy;
- secondary findings policy;
- reanalysis options.
Secondary findings are medically actionable pathogenic variants unrelated to the original reason for testing. Professional recommendations specify genes and variant types that may be deliberately assessed. Patients may be offered an opt-in or opt-out choice depending on local policy.
Incidental findings arise unexpectedly outside a planned list. Examples include evidence of consanguinity, a chromosome-sex difference, an unexpected family relationship, or a variant suggesting risk not covered by the requested analysis. Consent should address how such information is managed.
Privacy is especially important because a genome is identifying and contains information about relatives. Storage may include raw reads, aligned files, variant files, and reports. Patients can ask how long data and specimens are retained, whether data are shared for research, and how access is controlled. Complete deletion from distributed research datasets may not always be possible once data have been shared under consent.
Insurance and discrimination protections vary by jurisdiction and may not cover all products, such as life or long-term-care insurance. These issues can be discussed before elective testing.
Confirmation, Reanalysis, and Next Steps
Confirmation depends on variant quality and consequence. A well-validated laboratory may report high-confidence small variants without Sanger confirmation. Complex, low-level, or treatment-defining findings may need an orthogonal method.
Common follow-up methods include:
- Sanger sequencing for a small sequence variant or breakpoint;
- MLPA, digital PCR, or microarray for copy-number change;
- karyotype or FISH for chromosome context;
- repeat-primed PCR or Southern blot for an expansion;
- RNA sequencing for splicing or expression effects;
- long-read WGS for phase, repeats, pseudogenes, and complex structures;
- methylation analysis for imprinting or epigenetic disorders;
- deep sequencing of affected tissue for mosaicism.
After a diagnostic result, parental or family testing establishes inheritance and identifies at-risk relatives. A targeted assay is generally used rather than repeating WGS. Reproductive risk depends on whether the condition is dominant, recessive, X-linked, mitochondrial, mosaic, or structural.
After a negative result, multidisciplinary review can determine whether the phenotype suggests a missed class. One person may need muscle RNA sequencing; another may need optical genome mapping; another may need biochemical testing rather than more sequencing. The next step should follow the suspected mechanism.
Reanalysis is one of WGS’s major strengths. Stored data can be reprocessed with new callers, compared with updated references, and interpreted using newly established gene-disease relationships. Reanalysis can reveal variants that were present from the start but not recognizable or reportable.
A useful reanalysis request provides updated symptoms, new family members, pathology, imaging, biochemical results, and any candidate diagnoses. Laboratories differ in timing and fees. Some perform periodic automatic review; others require a clinician request.
Reanalysis is limited by the original data. A region with no usable reads cannot be recovered computationally. In that case, resequencing with a newer short-read platform, higher coverage, or long-read technology may be necessary.
Long-read genome sequencing is especially promising for previously negative cases, but its added yield depends on patient selection and the completeness of earlier testing. New technology can expose more variants than science can currently interpret, so functional studies and international case sharing remain essential.
Patients should retain the complete genome report and future amendments. Important details include genome build, transcript, exact variant, zygosity, laboratory, and date. These allow relatives to receive correct targeted testing and help clinicians compare future reclassifications.
WGS offers the broadest single DNA sequencing view currently available in many clinics. Its value comes not from the word “whole,” but from rigorous analysis across validated variant classes, careful phenotype matching, transparent limits, and a plan to revisit unresolved data.
References
- Toward clinical long-read genome sequencing for rare diseases — 2025 Perspective.
- The additional diagnostic yield of long-read sequencing in undiagnosed rare diseases: a systematic review — 2025 Systematic Review.
- Identification of technically challenging variants by whole-genome sequencing in rare disease diagnostics — 2025 Study.
- Genome Sequencing for Diagnosing Rare Diseases — 2024 Study.
- Lessons and pitfalls of whole genome sequencing — 2024 Review.
- Evidence from 2100 index cases supports genome sequencing as a first-tier genetic test — 2024 Study.
Disclaimer
This article provides general education and does not replace advice from a clinical genetics professional or testing laboratory. WGS platforms, analysis pipelines, reporting scope, and secondary-finding policies vary and continue to change. Medical decisions should rely on the complete report, appropriate confirmation, and the patient’s clinical and family context.





