AgSpec uses retrieval-based drafting to speed up coding agents by up to 4.76x

AgSpec uses retrieval-based drafting to speed up coding agents by up to 4.76x

A new paper presents AgSpec, a framework for retrieval-based speculative decoding (SD) in coding-agent pipelines. Retrieval-based SD drafts tokens by copying continuations from existing text. The authors say this suits coding agents, which repeatedly reproduce code, logs and earlier attempts.

The authors argue that existing retrieval methods fall short in agent pipelines for two reasons. First, much of the reusable text is missing from their corpora, or is stored in a form that differs from what the agent actually emits. Second, their draft lengths ignore the fact that accept length varies across agents and drifts over turns. AgSpec is described as supplying the corpus and draft-length policies that existing retrieval engines lack in these pipelines.

On the corpus side, AgSpec retrieves from session, workspace and global corpora. It retains the ongoing session trajectory and indexes opened files in the agent's emission format. On the draft-length side, it bounds each agent's draft length with an offline-profiled cap, then adapts the length online from verification feedback.

The reported results come from two repository-level multi-agent coding benchmarks. There, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, not all of them. Compared with plain autoregressive decoding, it raises generation throughput by up to 4.37x at batch size 1 and up to 4.76x at batch size 16. The authors add that AgSpec also stays effective on benchmarks without a repository or a multi-agent pipeline, which they say shows its gains generalize to coding agents broadly.

Key facts

  • AgSpec is a retrieval-based speculative decoding framework for coding-agent pipelines; it drafts tokens by copying continuations from existing text such as code, logs and earlier attempts.
  • It retrieves from session, workspace and global corpora, keeps the ongoing session trajectory, and indexes opened files in the agent's emission format.
  • Each agent's draft length is bounded by an offline-profiled cap and then adapted online from verification feedback.
  • On two repository-level multi-agent coding benchmarks it beats five retrieval-based drafters and EAGLE-3 in most evaluated settings.
  • Throughput gains over autoregressive decoding reach up to 4.37x at batch size 1 and 4.76x at batch size 16.

Why it matters

Coding agents spend much of their time reproducing text they have already seen: code, logs and earlier attempts. Retrieval-based speculative decoding exploits that by copying continuations instead of generating them from scratch. The paper's point is that generic retrieval engines miss reusable text and use draft lengths that do not fit how agents behave, and AgSpec targets exactly those two gaps. If the reported gains hold, agent pipelines could generate noticeably faster.

Who it affects

The work is aimed at people building or serving coding-agent pipelines, especially multi-agent setups that operate on a repository. It is also relevant to anyone comparing speculative decoding approaches such as retrieval-based drafters and EAGLE-3 for agent workloads.

How to use it

The source describes the design rather than a release. In practice the framework has two parts: corpus policy (session, workspace and global corpora, with opened files indexed in the agent's emission format) and draft-length policy (an offline-profiled cap per agent, adapted online from verification feedback). Anyone wanting to apply the idea would need to build or obtain an implementation. No code or release availability is mentioned in the source.

How solid is it

The claims come from the paper's abstract and are the authors' own. The headline numbers are maxima: up to 4.37x at batch size 1 and up to 4.76x at batch size 16 over autoregressive decoding. AgSpec outperforms the five retrieval-based drafters and EAGLE-3 in most evaluated settings, not every one. The benchmarks are not named. The models used, hardware and baseline drafter names (other than EAGLE-3) are not given. No average or median speedup is stated, and no figures are given for the generalization benchmarks.

Risks and caveats

Because the speedups are reported only as maxima, typical gains may be lower, and the source does not say which setting produced the best numbers. The advantage is claimed in most, not all, evaluated settings. The claim that gains generalize to coding agents broadly rests on benchmarks without a repository or a multi-agent pipeline, and no figures for those are given. Independent replication would be needed before relying on the numbers.