// NATURE NEWS — SPAZIO & SCIENZA
Retrofitted LLM can count the letter ‘i’s in ‘artificial intelligence’
Zhao Zhang is in the Key Lab of High Confidence Software Technologies, School of Computer Science, Peking University, Beijing 100871, China.
Search author on:
PubMed
Google Scholar
Yingfei Xiong is in the Key Lab of High Confidence Software Technologies, School of Computer Science, Peking University, Beijing 100871, China.
Search author on:
PubMed
Google Scholar
Strawberry contains the letter ‘r’ three times, but when asked, many large language models (LLMs) answer that the letter appears twice. This happens because most LLMs encode words as ‘tokens’ that represent sequences of letters. LLMs that operate in this way can achieve excellent performance, but they cannot access the individual characters in each word, which are encoded as binary sequences called bytes. Writing in Nature, Minixhofer et al.1 now report an approach called byteification that retrofits token-based LLMs to operate at the byte level. The authors show that byteified models can achieve competitive performance while retaining the ability to read individual characters.
Access Nature and 54 other Nature Portfolio journals
Get Nature+, our best-value online-access subscription
Prices may be subject to local taxes which are calculated during checkout
Minixhofer, B. et al. Nature https://doi.org/10.1038/s41586-026-11111-4 (2026).
Sennrich, R., Haddow, B. & Birch, A. in Proc. 54th Ann. Meet. Assoc. Comput. Linguist. 1715–1725 (2016).