// NATURE NEWS — SPAZIO & SCIENZA
Predicting genome-wide functional constraints with GPN-Star
Nature
(2026) Cite this article
Genomic language models have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences1. However, standard genomic language models adapted from natural language processing often require large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks2,3,4. Here we introduce a genomic pretrained network with species tree and alignment representations (GPN-Star), which is a biologically grounded genomic language model featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammal and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales show task-dependent advantages of modelling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms previous methods in prioritizing pathogenic and fine-mapped genome-wide association study variants, yields strong enrichments of complex trait heritability and improves power in rare variant association testing5. Extending beyond humans, we train GPN-Star for five model organisms—Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana—demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful and flexible tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.
A fundamental challenge in biology is understanding the functional significance of genetic variants. Despite tremendous advances in experimental technologies in genomics, determining which variants influence phenotypes or contribute to diseases remains difficult. Evolution provides valuable insights: deleterious mutations tend to be purged by natural selection, although tolerated or advantageous changes may accumulate. Millions of years of evolution therefore provide a genome-wide record of functional constraint for variant interpretation. This idea has deep roots in molecular evolution and comparative sequence analysis6. More recently, large comparative genomics studies have quantified evolutionary constraints across the genomes of hundreds of species7,8.
Genomic language models (gLMs) have emerged as a promising approach for extracting evolutionary information directly from raw DNA sequences through self-supervised learning (ref. 1 and references therein). By predicting masked nucleotides from sequence context, gLMs estimate per-site likelihoods that reflect evolutionary constraint without requiring labelled data. These likelihoods have been shown to be effective predictors of genome-wide variant effects1,9. However, gLMs based on standard language modelling frameworks still underperform compared with much simpler classical phylogenetic models on certain variant interpretation tasks—particularly in complex eukaryotic genomes such as in humans and in distal regulatory elements such as enhancers2—even with massive model sizes3,4.
A long-standing approach to modelling evolutionary data is to construct multiple sequence alignments (MSAs). By algorithmically aligning homologous loci across biological sequences, MSAs show site-specific patterns of conservation and facilitate the inference of evolutionary preferences for variants at specific positions. Protein MSAs have enabled highly successful models, including AlphaFold10, MSA Transformer11 and EVE12, and have recently experienced renewed interest as scaling single-sequence protein language models has shown diminishing returns13,14,15. Extending this concept from proteins to