// NATURE NEWS — SPAZIO & SCIENZA
Retrofitting language models to operate over bytes
Nature
(2026) Cite this article
Recent advances in artificial intelligence (AI) have largely been driven by large language models, deep neural networks that operate over discrete units called tokens. To represent text, most large language models use words or word fragments as the tokens, known as subword tokenization1. Subword tokenization obscures fine-grained information, which is problematic, especially for scientific data—such as computer code or biological sequences—where meaning depends on the individual characters or bytes2. Models that instead operate directly on the byte encoding of text avoid these limitations, but until now they have lagged behind subword-based models in performance. Here we introduce a general method for creating byte-level large language models through byteification that approach the capabilities of subword-based systems. We use a two-stage conversion procedure to retrofit existing subword-based models into byte-level models with minimal extra training. The resulting models outperform earlier byte-level approaches and excel on character-level reasoning tasks, achieving practical inference speeds by efficiently processing byte-level information and adaptability by reusing the existing ecosystem around the source large language model. Our results remove a long-standing performance barrier to end-to-end byte-level language modelling, demonstrating that models operating on raw text encodings can scale competitively while offering advantages in domains requiring fine-grained textual understanding.
Recent progress in AI has been driven by end-to-end deep learning systems that learn representations directly from data. Large language models (LLMs) exemplify this trend, achieving strong capabilities by training on massive collections of text3,4. However, despite their apparent generality, contemporary LLMs are not fully end-to-end: before learning can begin, text must first be mapped to a sequence of discrete units called tokens. The choice of tokens, although sometimes overlooked, fundamentally shapes the representations that LLMs learn and the behaviours that they exhibit5,6,7,8,9.
The vast majority of contemporary LLMs use words or parts of words as the tokens in a process known as subword tokenization1,10. This leads to many problems. LLMs that use subword tokenization suffer from limited character-level understanding11,12,13, which especially hinders performance with scientific data, such as code and biological sequences2,14,15,16,17; they are also implicitly biased towards generating particular responses based on how the prompt is tokenized18,19,20, they are restricted in the number of words they can incorporate in their vocabulary, which in practice leads to English-centricity6,21, and they potentially suboptimally allocate their compute2,22. These problems have motivated extensive research into alternatives to subword tokenization, most commonly by using the underlying UTF-8 bytes that the text is encoded as23 as the discrete units. Many earlier byte-level LLMs claim to outperform subword-level LLMs on the efficiency–performance Pareto frontier2,22,24,25,26,27. However, in practice, byte-level LLMs have not seen widespread adoption so far, and all leading LLMs still exclusively rely on subword tokenization.
We hypothesize that the key reason for this mismatch between theory and practice is that existing approaches to byte-level language modelling focus predominantly on training a new byte-level model from a random initialization and comparing it against a subword-level LLM also trained from a random initialization. By contrast, the training of state-of-the-art subword-level LLMs is rapidly evolving, combining innovations in training data curation, model architecture and post-training. Keeping up with this pace is infeasible for byte-level LLM development without extensive investments.
To resolve this mismatch, we introduce a general method for creating byte-level L