// NATURE NEWS — SPAZIO & SCIENZA
An Icelandic pangenome reference
Nature
(2026) Cite this article
Reference bias is an issue that affects most genomic studies analysing short reads mapped to a reference genome1,2. It can be mitigated by mapping to multiple haplotypes represented in a pangenome3,4,5. Here we introduce two new methods to address reference bias: Emblask for pangenome construction and Weaver for mapping to pangenomes at scale. Emblask is a hybrid long- and short-read haplotype-resolved dual assembly pipeline for parent–offspring trio data. Using Emblask, we assembled 698 Icelandic haplotypes and added them to the Human Pangenome Reference Consortium (HPRC) pangenome4 to construct an Icelandic pangenome reference (HPRC-ICE) including 51.41 million small variants. We mapped the short reads of 57,630 Icelanders to HPRC-ICE with Weaver and called 98.96 million variants, representing a 6.17% increase over mapping to a linear reference. We uncovered new variants in low-mappability regions, including a pathogenic single nucleotide polymorphism (SNP) in GBA1 that associates with early onset Parkinson’s disease and a missense SNP in CBS that is pathogenic for homocystinuria. We replicated the GBA1 association in the UK Biobank6 with a targeted remapping of 429,193 British and Irish participants.
Analyses of human genetic sequences7,8,9 are predominantly based on short reads mapped to a single reference genome10,11, resulting in an inaccurate or an incomplete mapping owing to gaps in the reference sequences and a strong bias towards the reference allele1,2. The current human reference genome11 (GRCh38) has been updated several times since its initial release with alternate contigs that represent chromosomal sequences diverging from its primary scaffolds. Although those additional sequences contribute to more diversity in GRCh38, they are often discarded from genomics analysis because of their similarity with primary scaffolds, which often leads to mapping ambiguity and in turn decreased variant calling sensitivity. Furthermore, GRCh38 is still an incomplete representation of the human genome, as its primary scaffolds contain about 151 Mb of gaps and 9 Mb of erroneously duplicated or collapsed regions12. The recent advance of long read sequencing13 has enabled the first gapless assembly of a human genome14, T2T-CHM13, comprising near error-free assemblies of the 22 autosomes and chromosome X, later complemented by a chromosome Y assembly15. Although T2T-CHM13 adds more than 182 Mb of sequence with respect to GRCh38, including 99 novel protein-coding genes, it does not capture any sequence differences between haplotypes or individuals.
Computational pangenomics is a field of study dedicated to the joint analysis and usage of genomic sequence collections over traditional linear references3,16,17,18,19,20,21,22. To this end, the HPRC has sequenced the genomes of 47 individuals from various ethnic groups and assembled them into contiguous high-quality haplotype-resolved assemblies4. The genetic variations of these individuals have been extracted from their assemblies and combined into a pangenome represented as a graph (Methods). Overall, the HPRC pangenome contributes an additional 90 Mb of non-reference sequences derived from structural variants and 119 Mb of euchromatic sequences with respect to GRCh38. Despite this, widespread adoption of pangenomes has been hindered due to their structural complexity, usage of new file formats and more computational requirements than required for linear reference mapping.
Here we used the Emblask pipeline to assemble 698 high-quality Icelandic haplotypes and combined them with the HPRC assemblies to create the HPRC-ICE pangenome reference, composed of 788 haplotypes (Fig. 1). Weaver maps to pangenome references composed of hundreds of haplotypes, such as HPRC-ICE, and outputs read mappings in a linear coordinate system such as GRCh38 or T2T-CHM13. As Weaver uses existing file formats, it can simply replace any linear refer