KVFetch adds position-based prefetching to KV cache compression
An arXiv paper argues that KV cache compression, now seen as essential as context windows grow to tens or hundreds of thousands of tokens, has a structural blind spot. The authors group existing methods into three families: score-based eviction, summary compensation, and offload-and-recall. All three, they say, decide what to keep or recall by content relevance to the current query.
Their claim is that a cache supports two access modes: associative lookup by content and sequential traversal by position. Current compressors implement only the first. That matters because retrieval-augmented generation, code completion and structured-data extraction all require the model to reproduce identifiers, field values or code tokens verbatim from the context. Under compression, content-based eviction keeps the head of such a sequence but discards its continuation, so verbatim copying breaks irreversibly midway. The authors call this failure sequential forgetting. They say it resists better scoring, larger budgets, summary compensation and dynamic re-scoring, and that it is the dominant source of remaining quality loss under compression.
The proposed fix is KVFetch, described as a training-free, drop-in framework that opens a temporal recall channel for any score-based compressor. It works in three steps: it demotes evicted candidates to a quantized cold tier, detects active copying through a monotone read pointer, and prefetches positional successors into fixed-size hot-tier slots, all without increasing attention cost.
On RULER-16K under an iso-budget control, the paper reports that KVFetch recovers verbatim copying from 0.8 to 78.4 and raises the 13-task average by +8.4, with gains concentrating on tasks that require sequential access. On LongBench, where no task requires sequential access, the channel stays dormant and imposes no cost.
Key facts
- The paper names three families of KV cache compression (score-based eviction, summary compensation, offload-and-recall) and says all three choose what to keep by content relevance to the current query.
- It identifies a failure called sequential forgetting: eviction keeps the head of a sequence but drops its continuation, so verbatim copying breaks midway.
- KVFetch is training-free and drop-in for any score-based compressor; it uses a quantized cold tier, a monotone read pointer and prefetching of positional successors into fixed-size hot-tier slots.
- On RULER-16K under an iso-budget control, verbatim copying goes from 0.8 to 78.4 and the 13-task average rises by +8.4.
- On LongBench, where no task requires sequential access, the added channel stays dormant and imposes no cost.
Why it matters
KV cache compression is how long-context inference is kept affordable, and this paper claims the whole field shares one design gap: it can find tokens by content but cannot walk forward by position. If the authors are right, sequential forgetting is the dominant source of remaining quality loss under compression, and tuning scoring or raising budgets will not remove it. The reported jump in verbatim copying on RULER-16K, from 0.8 to 78.4, is the headline evidence for that argument.
Who it affects
The paper points at workloads that must reproduce text exactly from the context: retrieval-augmented generation, code completion and structured-data extraction, where identifiers, field values or code tokens have to be copied verbatim. Anyone building or running compressed-cache inference for such tasks is the natural audience, along with researchers working on score-based eviction, summary compensation and offload-and-recall methods.
How to use it
The authors describe KVFetch as training-free and drop-in, a layer that can sit on top of any score-based compressor without retraining. No code release is mentioned in the text, so there is nothing concrete to install from this source alone. In practice, this is a design to evaluate on your own copying-heavy workloads once an implementation is available.
How solid is it
This is an arXiv preprint, and every result comes from the authors' own description. The main number is measured on RULER-16K under an iso-budget control, which puts the baseline and KVFetch on the same budget. The gains are said to concentrate on tasks requiring sequential access, which fits the proposed mechanism. The text names no authors or institutions, and it does not name the models, compression budgets or compressors used in the experiments.
Risks and caveats
The 0.8, 78.4 and +8.4 figures come with no stated unit, and the baseline value of the 13-task average is not given, so the relative size of the +8.4 gain is unknown. LongBench is reported only as a dormant channel with no cost, with no numbers. No memory, latency or throughput figures are given, so the real overhead of the quantized cold tier and prefetching is not visible. The claim that the channel adds no attention cost is the authors' own.
“We show this shared design is structurally incomplete.”
— KVFetch paper, arXiv 2610.08811