WhiteMatter lets every Transformer layer reuse any depth's KV cache

A standard Transformer produces representations of past tokens at every layer while generating text, but each layer can normally use only representations from its own depth. The paper argues this restriction stops the model from fully reusing information it has already computed.
The proposed fix is WhiteMatter, an architecture in which every layer can draw on past-token representations from any depth. A learned mixer picks the most useful depths for the current context and combines their representations into shared key-value (KV) cache channels. Because these channels are shared across layers, the cache can get smaller.
The reported results are relative. Given the same number of training tokens, WhiteMatter with a full-size cache performs comparably to a standard Transformer with 50% more layers. With half the KV cache, it outperforms matched standard Transformers at two model scales, up to 1.3B parameters.
There is a cost. Cross-layer connections introduce dependencies that slow both training and prompt processing. To address this, the paper proposes cyclic iteration, which updates interleaved groups of tokens in turn while processing the tokens within each group in parallel. On a reference model trained with exact autoregressive execution, cyclic iteration converges 12.5x faster than standard Jacobi iteration.
Key facts
- WhiteMatter lets every layer draw on past-token representations from any depth, instead of only its own depth as in a standard Transformer.
- A learned mixer selects the most useful depths for the current context and combines them into shared KV cache channels.
- With a full-size cache and the same training tokens, it performs comparably to a standard Transformer with 50% more layers.
- With half the KV cache, it outperforms matched standard Transformers at two model scales, up to 1.3B parameters.
- Cyclic iteration, which offsets the slowdown from cross-layer dependencies, converges 12.5x faster than standard Jacobi iteration on a reference model trained with exact autoregressive execution.
Why it matters
The KV cache is what a Transformer keeps about past tokens while generating text, and the paper's starting point is that layers cannot fully reuse what is already computed because each one sees only its own depth. WhiteMatter loosens that rule and reports two gains: quality comparable to a model with 50% more layers at a full-size cache, and better results than matched standard Transformers with half the cache.
Who it affects
The paper is aimed at people designing Transformer architectures and those who care about cache size and parameter efficiency. The results cover two model scales, up to 1.3B parameters, so they speak to small and mid-sized models rather than frontier systems.
How to use it
The source text describes the method only at the level of ideas: a learned mixer over depths feeding shared KV cache channels, plus cyclic iteration for training and prompt processing. It mentions no code or model release, so this is a design to study rather than something to run today.
How solid is it
The claims come from the authors' own summary. The comparisons are stated against matched standard Transformers, but the text names no benchmarks, datasets or metrics behind 'performs comparably' or 'outperforms'. The 12.5x figure applies to a reference model trained with exact autoregressive execution, not to all models or to the two tested scales.
Risks and caveats
Cross-layer connections add dependencies that slow training and prompt processing, and cyclic iteration is the proposed remedy. No absolute speed figures or wall-clock comparison with a standard Transformer are given, so the net cost is unclear. The cache saving is quantified only as half at the tested scales, and the statement that sharing channels 'can reduce' cache size is not itself a measured result.