Q&D trains agents to ask for information users never requested

A paper on Hugging Face studies what a tool-using agent should pursue when the user has not asked for it. The starting point is that an agent usually answers what the user explicitly requests, yet finishing the task may need information the user never mentioned. The authors say earlier work on proactive agents mainly asks whether and when an agent should act on its own, not what information it should go after. They study a separate axis: the content of proactivity.
They define two kinds. Horizontal proactivity pursues unstated information that the current context already identifies. Vertical proactivity pursues needs that only earlier evidence reveals. To measure both, they build a need graph recovered from a benchmark's own decomposition of a task. The graph records which needs depend on which, so both forms of proactivity, and whether the agent stops at the right time, can be scored from a transcript without a model judge.
To teach the behavior, the authors propose Q&D (questioner and drafter). It trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge.
The evaluation uses held-out splits of three multi-hop question-answering benchmarks. At equal retrieval spend, the trained questioner improves both horizontal and vertical proactivity over the same model when prompted. It also outperforms a prompted model 15 times larger in the same role on two of the three benchmarks. The gain persists after controlling for question volume and length.
The authors then test transfer. Without further training, they place the questioner in an interactive customer-service agent with a simulated customer. There it completes more tasks while asking fewer questions, and in retail it outperforms the 15 times larger model with fewer follow-up turns from the customer. Their conclusion: proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.
Key facts
- The paper separates horizontal proactivity (unstated information the current context already identifies) from vertical proactivity (needs that only earlier evidence reveals).
- A need graph recovered from a benchmark's own decomposition lets both forms, and whether the agent stops at the right time, be scored from a transcript without a model judge.
- Q&D (questioner and drafter) trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge.
- On held-out splits of three multi-hop QA benchmarks, at equal retrieval spend, the trained questioner beats the same model prompted and beats a prompted model 15 times larger on two of the three.
- With no further training, the questioner completes more tasks with fewer questions in a simulated customer-service setting, and in retail beats the 15 times larger model with fewer follow-up turns from the customer.
Why it matters
The paper shifts the question about proactive agents. Earlier work, as the authors describe it, focuses on whether and when an agent should act on its own. This one asks what the agent should go after, and when it should stop. The horizontal and vertical split gives a vocabulary for that: some missing information is visible from the current context, and some only becomes visible once earlier evidence has come back. The scoring method matters too, since the need graph allows evaluation from a transcript without a model judge.
Who it affects
Researchers building tool-using and question-asking agents are the main audience, along with teams working on multi-hop question answering. The customer-service result is relevant to anyone deploying agents that talk to users, because the reported gain is more completed tasks with fewer questions, and in retail fewer follow-up turns from the customer.
How to use it
The abstract describes a method, not a product. Q&D trains a questioner using the outcome of the continuation, meaning how much of the required evidence it retrieves, rather than a reward model or judge. The need graph gives an evaluation recipe that reuses a benchmark's own task decomposition. No code or data release is mentioned.
How solid is it
The claims come from the paper's abstract, and the authors report the gains at equal retrieval spend on held-out splits of three benchmarks. They also say the gain persists after controlling for question volume and length, which addresses the obvious objection that the trained model simply asks more. The larger-model comparison is weaker than it may sound: the trained questioner beats the 15 times larger prompted model on two of the three benchmarks, not all three. The customer-service test uses a simulated customer.
Risks and caveats
No numerical results are given, only qualitative improvements, so the size of the gains is unknown. The names of the benchmarks and models are not given, and the abstract does not say which two of the three benchmarks the larger-model win covers. The size of the gain in the customer-service and retail settings is not quantified, and that test relies on a simulated customer rather than real users.
“These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.”
— From the paper's abstract