Byteification retrofits subword LLMs into byte-level models

Byteification retrofits subword LLMs into byte-level models

Almost every leading large language model reads text as subword tokens, meaning words or word fragments. The authors of a paper on nature.com argue that this hides fine-grained information, which hurts most on scientific data such as computer code or biological sequences, where meaning depends on individual characters or bytes. Models that work directly on the byte encoding of text avoid that problem, but until now they have lagged behind subword models in performance, and none of the leading LLMs uses them.

The authors hypothesize that earlier byte-level work went wrong by training a new model from a random initialization and comparing it with a subword model trained the same way. Subword training keeps improving in data curation, architecture and post-training, and keeping pace is infeasible for byte-level development without extensive investment. Their answer is to start from an existing subword model instead. They call the process byteification: a two-stage conversion procedure that retrofits a subword LLM into a byte-level one with minimal extra training, using less than 1% of a typical pretraining budget (49.1B tokens in total).

The paper applies it to four models. Bolmo 7B and Bolmo 1B come from the fully open Olmo 3 7B and OLMo 2 1B. Bwen 8B comes from Qwen3 8B Base, and Blama 8B from Llama 3 8B.

The architecture first aggregates the byte stream into patches of one or more bytes, runs the patches through a large transformer, and then depools them back into bytes. The authors call this style a latent tokenizer language model (LTLM); DTP, BLT and H-Net are earlier examples. A shallow but wide local encoder builds byte representations, a boundary predictor decides where patches start and end, a pooling module reduces each patch to one representation, a deep global model processes the patches, and a local decoder predicts the next byte. The local layers are matrix long short-term memory (mLSTM) layers, chosen because they kept high performance at high throughput in the authors' experiments.

The authors say their primary departure from earlier architectures is non-causal patch boundary prediction. Earlier boundary predictors look only at past context. Subword tokenizers, though, use information about future bytes when placing a token boundary: in their example, whether '_Wor' becomes a token depends on whether the text continues as '_Wor!' or '_World!'. The new predictor uses up to 1 byte of future context, which the authors found largely sufficient to match the behaviour of subword tokenization.

On results, the byteified models outperformed, on average, all earlier publicly available byte-level LLMs of comparable size. Bolmo 7B achieved a +16.5% absolute improvement in STEM tasks over BLT 7B, which was trained from random initialization. Bolmo 7B also greatly outperformed its source Olmo 3 on character understanding and was better in certain coding settings. Bwen 8B outperformed Bolmo 7B and reached performance close to, and sometimes surpassing, the source Qwen model. The authors add that byteified models can be sped up further by training with higher ratios of bytes per patch, and that existing components of the source model's ecosystem can adapt a byteified model without extra training cost.

The authors conclude that byte-level LLMs offer substantial promise as a foundation for future language models, with potential advantages in computational efficiency (lower energy and deployment costs), in reducing biases introduced by English-centric subword tokenization, and in applications that need fine-grained textual understanding, particularly in scientific and technical domains.

Key facts

  • Byteification is a two-stage procedure that retrofits an existing subword LLM into a byte-level model using less than 1% of a typical pretraining budget (49.1B tokens in total).
  • The paper produces Bolmo 7B and Bolmo 1B (from Olmo 3 7B and OLMo 2 1B), Bwen 8B (from Qwen3 8B Base) and Blama 8B (from Llama 3 8B).
  • Bolmo 7B scored a +16.5% absolute improvement in STEM tasks over BLT 7B, which was trained from random initialization.
  • Bwen 8B outperformed Bolmo 7B and reached performance close to, and sometimes surpassing, the source Qwen model.
  • The main architectural change is a non-causal boundary predictor that uses up to 1 byte of future context, which the authors found largely sufficient to match subword tokenization.

Why it matters

Byte-level models have long promised better character-level understanding and less dependence on a fixed vocabulary, but the authors say they have not been adopted and all leading LLMs still rely on subword tokenization. This paper's pitch is that the obstacle was cost: nobody could keep a from-scratch byte-level model current with fast-moving subword training. Converting an existing model for under 1% of a typical pretraining budget changes that equation, and the authors say it removes a long-standing performance barrier to end-to-end byte-level language modelling.

Who it affects

Researchers building or studying tokenizer-free and byte-level language models are the most direct audience, since the authors argue that cheap byteification can quickly surface promising architectures to train from random initialization. The paper singles out scientific and technical domains such as code and biological sequences, where character-level detail matters. The authors also point to users of non-English text, citing the English-centric bias that subword vocabularies introduce.

How to use it

The recipe starts from an existing open subword model and applies the two-stage conversion with a patch-based byte architecture; the paper demonstrates it on Olmo 3 7B, OLMo 2 1B, Qwen3 8B Base and Llama 3 8B. The authors say components from the source model's ecosystem can be used to adapt a byteified model with no extra training cost. The visible text carries no statement about public release or licences of the Bolmo, Bwen or Blama models.

How solid is it

The claims here are the authors' own, and the comparisons are described as averages across earlier publicly available byte-level models of comparable size. The wording on the source models is measured: Bwen 8B reached performance 'close to and sometimes surpassing' Qwen, and Bolmo 7B was better only 'in certain coding settings'. The benchmark and metric behind the +16.5% STEM figure are not named in the text available for this retelling, and the methods and full results tables were not available either, so the detail rests on the abstract and the opening of the main text.

Risks and caveats

The authors present the reason earlier byte-level models lagged as a hypothesis, not a proven cause. Efficiency gains in energy and deployment cost are described only as potential advantages, and no measured savings are reported in the text available. No numerical results appear for Blama 8B or Bolmo 1B, and no inference-speed figures are given, although the abstract claims practical inference speeds. Byte-level models mostly land close to their subword sources and only sometimes surpass them, so this is a route to near-parity, not a clear win.

“Our results remove a long-standing performance barrier to end-to-end byte-level language modelling”

— the paper's abstract