For generations, the map of the human genome has been drawn from a single vantage point — one reference, one haplotype, one inherited set of assumptions. A team of researchers has now charted both copies of a human genome, maternal and paternal, with near-perfect fidelity, creating a benchmark called T2T-HG002 that challenges the foundational logic of how genetic medicine reads the book of life. The work, now enshrined as a National Institute of Standards and Technology reference material, does not merely improve an old tool — it proposes a different kind of tool entirely, one that begins with
Near-perfect diploid genome benchmark advances personalized medicine
Complete genome assembly recovered two to seven percent more sequence than variant calling
Why does it matter that this benchmark captures both maternal and paternal copies instead of just one?
Because most human genetic diseases are inherited from one parent or the other, and some genes exist in only one parental copy. If you only have a haploid reference, you miss those variants entirely. You're comparing apples to a single reference apple when you should be comparing apples to both the maternal and paternal apples.
So the old way was missing genes?
Not missing them entirely, but missing the variation in them. A gene might exist in a patient's maternal copy but not the paternal copy, or vice versa. The old reference couldn't see that difference because it only represented one haplotype. You'd sequence a patient and call variants against the reference, but you'd never know if a gene was absent or just different.
What about the 918 megabases of new sequence? That sounds like a lot.
It is. That's about 11.7% of the autosomal genome that earlier benchmarks simply didn't include. These are mostly repetitive regions, duplicated genes, and satellite DNA—the hard stuff that short-read sequencing can't handle. But these regions matter clinically. Immune genes, copy-number variable genes, disease-associated loci. They were dark matter in the old reference.
Why was it so hard to sequence these regions before?
Repetitive DNA looks identical to itself. If you have a thousand copies of the same sequence, short reads can't tell you which copy they came from. Long-read sequencing solves that by reading through the repetition in one continuous stretch. But it took multiple technologies—PacBio, Oxford Nanopore, Hi-C scaffolding, even fluorescence microscopy—to get it right.
And the quality score went from Q63 to Q69. What does that mean in practice?
It means the error rate dropped from about one error per billion base pairs to one error per 68 billion. For a three-billion-base-pair genome, that's the difference between three errors and 0.04 errors. Essentially error-free for clinical purposes.
So what changes in the clinic?
Right now, clinicians call variants against a reference. With T2T-HG002, they could build a complete personalized genome for each patient and compare it to this benchmark instead. You'd see structural variants, copy-number changes, and haplotype-specific variants that variant calling misses. It's a different way of thinking about what a genome is.
The Pulse
- Decades of genomic medicine have rested on a reference genome that is incomplete and biased, leaving disease-causing variants hidden in the very regions most relevant to human health.
- Repetitive DNA, duplicated genes, and structurally complex chromosomal zones have consistently defeated short-read sequencing, creating blind spots that conventional benchmarks could not even measure.
- Using layered long-read technologies, three-dimensional scaffolding, and iterative error correction, researchers assembled both parental chromosome sets for a single individual at a quality of fewer than one error per 68 billion base pairs.
- The new benchmark unlocks 918 megabases of previously unmapped sequence — including genes tied to immune function that exist in only one parental copy — and reduced discrepancies with the older standard from nearly 6,000 variants to just 219.
- A companion software tool now allows any genome assembly to be evaluated against this diploid standard, and early results suggest that building genomes from scratch outperforms variant-calling by a factor of ten in accuracy.
- Ribosomal arrays remain unfinished, clinical translation pathways are still undefined, and the benchmark does not yet reflect the mutational complexity of tumor genomes — but the framework for truly personalized genomic medicine is now visible.
For generations, the map of the human genome has been drawn from a single vantage point — one reference, one haplotype, one inherited set of assumptions. A team of researchers has now charted both copies of a human genome, maternal and paternal, with near-perfect fidelity, creating a benchmark called T2T-HG002 that challenges the foundational logic of how genetic medicine reads the book of life. The work, now enshrined as a National Institute of Standards and Technology reference material, does not merely improve an old tool — it proposes a different kind of tool entirely, one that begins with the individual rather than the average.
For decades, genomic science has navigated by a single fixed star — a reference genome that is, by design, incomplete. It represents one haplotype, contains gaps in medically critical regions, and forces every patient's unique biology through a lens built from someone else's chromosomes. A research team has now built a fundamentally different kind of map: T2T-HG002, a near-complete diploid human genome benchmark capturing both the maternal and paternal copies of every chromosome, accurate to 99.4% across its full length.
The construction required layering multiple technologies — PacBio HiFi for raw sequence, Oxford Nanopore for additional coverage, Illumina reads for phasing, Hi-C and Strand-seq for three-dimensional scaffolding, and fluorescence microscopy to validate the hardest regions. After successive rounds of correction, the assembly reached a quality score of Q68.9, meaning fewer than one error per 68 billion base pairs. The result is not a haploid abstraction but a true diploid genome, including both sex chromosomes.
The benchmark adds 918.2 megabases of high-confidence sequence that earlier references simply lacked — among them genes like DUSP22 and complement factor H-related variants that affect immune function and exist in only one parental copy. Thirteen maternal-only and twelve paternal-only autosomal genes were identified that would be invisible in a conventional reference. Within shared regions, discrepancies with the older Genome in a Bottle benchmark fell from nearly 6,000 variants to 219, implying that most of the old differences were errors in the older standard itself.
Equally significant is the framework the team built around the benchmark. Their software tool GQC can evaluate any genome assembly against the diploid standard at the base, haplotype, and structural levels. Applied to five years of HG002 assemblies, it found that building genomes from scratch — rather than calling variants against a reference — recovered two to seven percent more sequence and achieved roughly tenfold greater accuracy. The implication is that the future of clinical genomics may lie not in variant calling at all, but in complete assembly for each patient.
Now designated a National Institute of Standards and Technology reference material, T2T-HG002 can serve as a universal yardstick for sequencing platforms and clinical methods alike. Gaps remain — ribosomal DNA arrays are unfinished, clinical translation protocols do not yet exist, and the benchmark does not capture the mutational complexity of tumor genomes. But the architecture is in place, and the path toward genomic medicine that begins with the individual rather than the average has, for the first time, become clearly visible.
For decades, scientists have relied on a single reference genome to understand human genetic variation—a shortcut that has worked well enough for many purposes but leaves blind spots in the most medically important places. A team of researchers has now built something fundamentally different: a complete map of both copies of a human genome, maternal and paternal, accurate to 99.4% across its entire length. The benchmark, called T2T-HG002, represents a shift in how genomics might work in the clinic, moving away from comparing individual genomes against a fixed standard and toward building personalized maps from the ground up.
The problem with conventional genome sequencing is structural. When researchers want to find disease-causing mutations, they typically take a patient's genetic data and compare it to an existing reference sequence. But that reference is itself incomplete and biased—it represents only one haplotype, one set of inherited chromosomes, and it contains gaps and errors in the regions that matter most for human health. Short-read sequencing technology struggles with repetitive DNA, duplicated genes, and the complex structural variations that cluster in these hard-to-sequence zones. Researchers end up missing variants in places like immunoglobulin loci and copy-number variable genes, the very regions where genetic disease often hides.
T2T-HG002 was built using a different architecture. Instead of starting with an existing reference and calling variants against it, the researchers sequenced a single individual—HG002, a well-studied cell line—using multiple long-read technologies: PacBio HiFi for raw sequence data, Oxford Nanopore for additional coverage, and supplementary Illumina reads for phasing. They used Hi-C and Strand-seq data to scaffold the assembly correctly in three-dimensional space, and fluorescence microscopy to validate the most difficult regions. Over successive rounds of error correction and validation, they refined the assembly until it reached a quality score of Q68.9, meaning fewer than one error per 68 billion base pairs. The result contains both the maternal and paternal copies of every chromosome, including the sex chromosomes—a true diploid genome, not a haploid abstraction.
The benchmark adds 918.2 megabases of high-confidence sequence that earlier reference genomes simply lacked: 701.4 megabases of autosomal DNA and 216.8 megabases of sex chromosome sequences. These are not trivial additions. They include entire genes that exist in only one parental copy, like DUSP22 and the complement factor H-related genes, which are known to vary between individuals and affect immune function. The team identified 13 maternal-only and 12 paternal-only autosomal genes—variants that would be invisible in a haploid reference. Within the regions both benchmarks cover, the new standard reduced discrepancies with the older Genome in a Bottle benchmark from nearly 6,000 variants down to 219, suggesting that most of the remaining differences are errors in the older benchmark itself.
What makes T2T-HG002 transformative is not just its accuracy but its framework. The researchers developed software called GQC that can evaluate any new genome assembly by comparing it directly to the diploid benchmark, distinguishing between the two parental copies and assessing accuracy at the base level, the haplotype level, and across large structural changes. When they applied this tool to five years of HG002 assemblies generated with different technologies, they found that de novo genome reconstruction—building a complete genome from scratch rather than calling variants against a reference—recovered two to seven percent more sequence and achieved roughly tenfold greater accuracy. This suggests that the future of personalized genomics may not be variant calling at all, but rather complete genome assembly for each patient.
The benchmark is now a National Institute of Standards and Technology reference material, meaning it can serve as a standard for evaluating sequencing platforms, assembly algorithms, and clinical interpretation methods. But significant work remains. Some regions, particularly the ribosomal DNA arrays, remain unfinished even in this near-complete benchmark. Clinical genomics still relies primarily on variant calling, and no standardized methods yet exist for translating whole-genome benchmarking performance into actionable clinical interpretation. The authors note that HG002 cells contain only low levels of somatic variation, so the benchmark may not fully capture the complexity of tumor genomes or tissues with high mutation rates. Still, the framework is in place. As sequencing technology continues to improve and researchers apply the same approach to diverse populations, the path toward truly personalized genomic medicine—one that does not force every patient's genome through the lens of a single reference—has become visible.
Notable Quotes
Unlike conventional human reference sequences, which generally represent a single haplotype, T2T-HG002 contains both maternal and paternal haplotypes, providing a diploid, two-haplotype representation.— Study authors, published in Cell