VGBench tests whether audio LLMs stay silent when a bystander speaks

VGBench tests whether audio LLMs stay silent when a bystander speaks

Audio language models can recognize spoken commands and call tools. The authors of a new paper argue that an agent has a prior job: it must first decide whether the acoustic and conversational context warrants action at all. To test that, they introduce VGBench, a 1,018-item diagnostic benchmark for what they call action-level addressedness. It covers three kinds of scenario: side-talk, self-talk and speaker-switch.

Every item uses the same action space of three options: silence, a tool call, or a natural-language answer. In the speaker-switch pairs, the specified words are held fixed while the source, the distance rendering and a temporal boundary define a controlled shift from the wearer to a bystander. So the same words should trigger action when the wearer says them and silence when someone else does.

The authors evaluated six raw Audio LLMs and three training-free adaptations. These systems often identify the target tool, yet rarely withhold action under the wearer-to-bystander shift. The highest raw switch mute rate was 14%.

The paper then uses VoxGate as a post-training case study. Supervised training mutes 91.3% of switched commands while choosing the correct tool for all nearby wearer commands and text-only controls. An exploratory GRPO stage has similar switch performance. In that stage, side-talk accuracy rises from 68.4% to 70.9%, and self-talk muting from 52.0% to 60.0%.

Factorized controls identify an independent source-change effect, while sensitivity to the far-field manipulation varies across acoustic renderings. The authors conclude that the benchmark measures multi-cue acoustic-context gating rather than isolated speaker identity.

Key facts

  • VGBench is a 1,018-item diagnostic benchmark covering side-talk, self-talk and speaker-switch scenarios, with a shared action space of silence, a tool call, or a natural-language answer.
  • Six raw Audio LLMs and three training-free adaptations often identify the target tool but rarely withhold action when the speaker shifts from wearer to bystander; the highest raw switch mute rate is 14%.
  • In the VoxGate case study, supervised training mutes 91.3% of switched commands while choosing the correct tool for all nearby wearer commands and text-only controls.
  • An exploratory GRPO stage has similar switch performance; side-talk accuracy rises from 68.4% to 70.9% and self-talk muting from 52.0% to 60.0%.
  • Factorized controls show an independent source-change effect, and the authors say the benchmark measures multi-cue acoustic-context gating, not isolated speaker identity.

Why it matters

Voice agents that call tools need more than speech recognition and tool selection. They need to know when a spoken command is meant for them. The paper's headline result is a gap: the tested models often pick the right tool yet rarely stay silent when the speaker changes from the wearer to a bystander, and the best raw switch mute rate is only 14%. VGBench gives a way to measure that failure directly, using one fixed set of words that should lead to action from one speaker and to silence from another.

Who it affects

The setup is framed around a wearer and nearby bystanders, so the findings bear most directly on people building or using voice agents that listen continuously in shared spaces. Researchers working on Audio LLMs and on post-training methods for them are the other audience: the paper offers both a diagnostic and a case study in fixing the gap.

How to use it

The abstract describes VGBench as a diagnostic: each item has three possible outputs (silence, a tool call, a natural-language answer), so a model can be scored on whether it withholds action as well as on whether it picks the right tool. The VoxGate case study shows one route to improvement, supervised post-training, with an exploratory GRPO stage on top. No code, dataset release, or availability is mentioned.

How solid is it

The numbers are specific and come with a clear design: a 1,018-item benchmark, six raw models and three training-free adaptations, and factorized controls that isolate a source-change effect. The abstract names no authors or institutions and does not name which Audio LLMs were tested. It does not say which base model VoxGate is built on or give the training data size. The GRPO stage is described by the authors as exploratory. The side-talk and self-talk figures are not explicitly labelled as supervised versus GRPO beyond the 'rises from' wording.

Risks and caveats

The 14% figure is the highest raw switch mute rate, not a typical one. The 91.3% result is for the VoxGate case study after supervised training, not for the raw models. Sensitivity to the far-field manipulation varies across acoustic renderings, so results may depend on how distance is simulated. The authors say the benchmark measures multi-cue acoustic-context gating rather than isolated speaker identity, so it should not be read as a speaker-recognition test. No deployment or real-world product results are reported.

“The benchmark therefore measures multi-cue acoustic-context gating rather than isolated speaker identity.”

— VGBench paper abstract