Paper predicts encoder-free multimodal LLMs catch up at about 10^22 FLOPs

Most modern multimodal large language models (MLLMs) are built on a pretrained visual encoder, which gives the model a strong visual prior. Encoder-free MLLMs take a different route: they learn visual representations directly from raw pixels, which gives a simpler and more unified architecture. Until now, the authors say, the scaling behavior of such models had not been systematically characterized. This paper fills that gap by comparing scaling laws for encoder-free and encoder-based MLLMs, and it reports three main findings.
First, removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models. The compute-optimal allocation for text stays nearly unchanged.
Second, the two architectures have nearly overlapping loss-compute frontiers on the text objective, but they diverge on the multimodal objective. Encoder-free models underperform at small scales, yet they are predicted to catch up at around 10^22 FLOPs, which the authors describe as well within practical pretraining budgets. This is a prediction from the scaling laws, not an observed training run.
Third, without a visual encoder the language model learns to take over the encoder's role through vision-specific adaptation. Three signs of this are reported: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated.
Overall, the authors conclude that the advantage of the visual prior provided by a pretrained encoder diminishes with scale. They position encoder-free architectures as a promising direction for multimodal pretraining.
Key facts
- The paper compares scaling laws of encoder-free MLLMs (which learn from raw pixels) with encoder-based MLLMs (which use a pretrained visual encoder).
- Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while the allocation for text is nearly unchanged.
- Loss-compute frontiers nearly overlap on text but diverge on the multimodal objective; encoder-free models lag at small scales and are predicted to catch up at around 10^22 FLOPs.
- Without an encoder, the language model adapts to vision: bidirectional interaction among visual tokens helps more with compute, visual processing moves to earlier layers, and expert routing for visual tokens concentrates.
- The authors conclude that the benefit of a pretrained encoder's visual prior diminishes with scale.
Why it matters
A pretrained visual encoder is the default building block of most modern multimodal models. Encoder-free designs are simpler and more unified, but nobody had systematically measured how they scale. This paper offers that measurement and argues that the encoder's advantage shrinks as compute grows. If the predicted catch-up at around 10^22 FLOPs holds, the extra encoder component would stop paying for itself at scales the authors call practical.
Who it affects
Teams that design and pretrain multimodal LLMs are the direct audience, especially those choosing between bolting a pretrained encoder onto a language model and training on raw pixels from the start. The compute-allocation finding also matters to anyone planning a pretraining budget, since the multimodal objective favors larger models once the encoder is removed.
How to use it
The abstract gives no code, models or data to reuse. The practical reading is as a planning input: the findings suggest that at small compute budgets an encoder-based design is the safer choice, while budgets near or above 10^22 FLOPs make an encoder-free design worth considering. Anyone acting on this should check the full paper's fitted laws against their own setup.
How solid is it
The source is a paper abstract, and the evidence behind the claims is not shown in it. The abstract names no authors or institutions. It gives no model sizes, parameter counts, datasets or benchmark scores. The key number, catch-up at around 10^22 FLOPs, is a prediction extrapolated from scaling laws, not an observed training run. The conclusions are the authors' own.
Risks and caveats
Scaling-law extrapolations can miss changes in behavior outside the range that was measured, so the catch-up point should be read as an estimate. The abstract does not say what size of model the shift toward larger models implies, or by how much. The catch-up is stated for the multimodal objective; on text the two architectures are nearly identical, so the result says nothing about a text-only advantage for either design.
“our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining”
— Paper abstract, How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining