// NATURE NEWS — SPAZIO & SCIENZA
Ultrafast and reference-free sequence discovery in single-cell data
Nature
(2026) Cite this article
Knowledge of RNA sequences, expression, splicing, isoforms, structure and modifications is central for understanding and targeting cellular processes. Revolutionary single-cell and spatial transcriptomics technologies—for example, as deployed by consortia such as the Human Cell Atlas—partially capture this diversity and generate cellular profiles that expand at petabyte scale each year1,2,3,4,5. Yet researchers cannot search sequences across these datasets: standard pipelines do not scale or rely on references, retaining only gene or isoform counts, whereas accessing raw sequences requires collecting, downloading and processing millions of large files. Here we present Malva, a computational platform that enables ultrafast, species-agnostic and reference-free interrogation of the raw sequence space, enabling searching for any sequence, mutation, splice junction or pathogen, or spatial location of arbitrary transcripts. The continuously expanding Malva Index currently comprises around 74 million cells from thousands of experiments in health and disease. Malva enables reference-free discovery—researchers can, for example, identify cell types and predict cell–cell similarity directly from sequence composition. Building on Malva’s speed and accuracy, we demonstrate how Malva can be flexibly connected to state-of-the-art neural networks and how to execute complex searches and enable automated analyses. Malva transforms single-cell atlases from static gene count tables into dynamic, sequence-resolved resources that may help to bridge human–machine reasoning about biology.
RNA is of fundamental importance for life. Knowledge of RNA sequences, RNA expression, splicing, isoforms, structure and modifications are central for understanding and targeting cellular processes. In the past decade, new sequencing technologies have made possible RNA sequencing (RNA-seq) in single, dissociated cells and subsequently for cells within whole tissue slices. Since cells are the basic units of life, these technologies have transformed life sciences, including medical and clinical research and applications, and annually generate petabytes of sequence information from hundreds of millions of cells1,6,7,8,9. The Human Cell Atlas consortium systematically quantifies RNA in healthy human cells2,3,10, for instance, while researchers also apply these technologies to diseases11. Moreover, vast amounts of data are generated for model systems, including organoids, and will be applied to investigate biodiversity on Earth as a whole12,13.
It is therefore of utmost importance to have tools that can both integrate these fragmented datasets into a unified resource and make them efficiently interrogable. Thirty-five years ago, BLAST pioneered this principle for sequence databases, transforming biological research by making it possible to locate a nucleotide sequence of interest in a vast collection of genomic sequences—critically, not only with accuracy but also with unprecedented efficiency14. These design principles made it broadly usable and ultimately indispensable across the life sciences. Yet, no existing computational tool—such as cell atlases (CELLxGENE, STOmicsDB, Single Cell Expression Atlas and others15,16,17,18,19,20,21)—can empower researchers to efficiently search for specific RNA sequences across the recently created vast single-cell, single-nuclei or spatial transcriptomics (SC/ST) data. At best, these platforms provide pre-computed gene counts that reflect the aggregated number of RNA sequences mapped to a reference. Therefore, critical information is missing: about transcripts that do not map to the reference (for example, transcripts generated by genetic lesions in cancer cells or transcripts that do not map to the reference haplotypes), RNA sequence editing, mutations, pathogens present in the cells and more.
Standard single-cell and single-nuclei pipelines remain reference-centric: th