HeteroFold moves KV cache between model families without prefill

HeteroFold moves KV cache between model families without prefill

Multi-agent LLM systems increasingly mix different models for specialized agent roles. The usual way for agents to talk is text, and that has a cost: each receiving model must prefill shared context that the sender has already processed. Reusing the sender's key-value (KV) cache would avoid the repeated work, but across model families the transfer has to cope with differences in tokenization, model depth and KV representations.\n\nThe authors propose HeteroFold, a prefill-free cross-family KV cache transfer method that keeps both the sender and the receiver frozen. It works in three steps: it aligns the model structures, maps the sender's cache into the receiver's space, and calibrates the result so the receiver's behavior is preserved.\n\nThe reported results cover six transfer directions. HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and on most short-context settings. On the multi-agent benchmark it matches text-based communication. For speed, the authors give one named pairing: at 32K context length, transfer from Llama-3.1-8B to Ministral-3-14B is 10.7 times faster than Native Prefill, and 1.18 to 1.47 times faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. The authors conclude that HeteroFold enables efficient cross-family KV reuse without receiver prefill.

Key facts

  • HeteroFold is a prefill-free method for transferring a sender model's KV cache to a receiver from a different model family, with both models kept frozen.
  • It aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior.
  • Across six transfer directions it gives the best cache-transfer performance on all four long-context benchmarks and most short-context settings.
  • On the multi-agent benchmark it matches text-based communication rather than beating it.
  • At 32K context, Llama-3.1-8B to Ministral-3-14B transfer is 10.7 times faster than Native Prefill and 1.18 to 1.47 times faster than Dense Latent and KV Ridge.

Why it matters

In multi-agent setups that combine different models, passing context as text forces every receiver to prefill material the sender has already processed. Sharing the KV cache directly removes that repeat work, but until the models share a family the caches do not line up: tokenization, depth and KV representations all differ. HeteroFold targets exactly that gap, and does so without touching the weights of either model. The headline figure is a 10.7 times speedup over Native Prefill at 32K context for one model pairing.

Who it affects

The paper speaks to builders of multi-agent LLM systems that assign different models to specialized roles, and to researchers working on KV cache reuse and prefill-free transfer. The only named pairing is Llama-3.1-8B to Ministral-3-14B, so the clearest relevance is to setups mixing models of that kind.

How to use it

This is a research result, not a product. The abstract describes the method as three steps (align model structures, map the sender cache into the receiver space, calibrate to preserve receiver behavior) and says both models stay frozen, so nothing is retrained in either model. It does not say whether code or weights are released.

How solid is it

The claims come from the authors' own abstract. They report results across six transfer directions, with the best cache-transfer performance on all four long-context benchmarks and most short-context settings, and parity with text-based communication on the multi-agent benchmark. The speed figures (10.7 times and 1.18 to 1.47 times) are stated for the Llama-3.1-8B to Ministral-3-14B direction at 32K context, not for every direction or length. No absolute latency or accuracy figures are given, only speed ratios at 32K context. The abstract names no authors or institutions.

Risks and caveats

HeteroFold matches text-based communication on the multi-agent benchmark; it is not reported as better. It is best on most, not all, short-context settings. The speedups apply to one direction at 32K context. The abstract does not report any cost of the method, such as calibration or training cost, and it does not name the other five transfer directions, the benchmarks or the multi-agent benchmark, so how far the results generalize cannot be judged from it.

“These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.”

— Abstract of the HeteroFold paper