For decades, biologists worked with one human reference genome: a consensus sequence assembled from a handful of donors and treated as the universal template. Every new genome we sequenced was compared against it — variants were whatever deviated from this single standard.
The problem is that no two humans are identical. Millions of short substitutions, hundreds of thousands of insertions, deletions, and structural rearrangements separate any two people. The reference captures the most common version at each position, but that means the reference is wrong for everyone in slightly different ways. Reads that span a region absent from the reference simply fail to map — a phenomenon called reference bias.
A pangenome graph replaces the linear reference with a directed graph. Sequence that is shared across all genomes becomes shared nodes; places where individuals differ branch into parallel bubbles — one path per variant. Every genome in the collection is represented as a walk through this graph. Nothing is discarded, nothing is privileged as the canonical sequence.
The Human Pangenome Reference Consortium published the first draft human pangenome in 2023, encoding 47 diverse genomes and capturing over 119 million base pairs of sequence not present in the original GRCh38 reference. The field is active and open — see related work on genome assembly and sequence alignment.
Comments
Loading comments...