SECRET: training-free method curbs audio-visual LLM hallucinations

Audio-visual large language models (AVLLMs) combine visual, auditory and linguistic information, and the paper's abstract says they have made remarkable progress in multimodal understanding and reasoning. But the authors point to a critical weakness reported in recent studies: source-confused grounding hallucination. This is when cues from the unused modality push the model toward a response that the required modality does not support. The authors say this undermines reliability in real-world applications.
Existing methods have made progress against the problem, but the authors say how the failure arises from internal cross-modal interactions remains insufficiently understood. To fill that gap, they run path-intervention and representation analyses. These reveal what they call a question-relay mechanism: the model's question states carry interfering cues alongside the evidence from the required source, and that weakens grounding in the required modality. One intervention result backs this up. Cutting the pathways from the interfering modality to the question states recovers more correct-answer logit than cutting the pathways to the generation position.
Building on this, the authors propose SECRET (SourcE-Conditioned RElay sTeering), a training-free method that reduces cross-modal interference at the question relay. It takes contrasting question representations, produced by applying different modality-pathway interventions, and uses them to steer the original question states toward the evidence from the required source.
The authors test SECRET on two widely used benchmarks, CMM and AVHBench, across three AVLLMs. They report that it consistently outperforms prior training-free methods and substantially mitigates source-confused grounding hallucinations, with gains of up to +18.0 and +7.1 percentage points over base models. Modality-specific captioning, they add, shows the method also generalises to open-ended generation.
Key facts
- Source-confused grounding hallucination in AVLLMs means cues from the unused modality induce responses that the required modality does not support.
- Path-intervention and representation analyses point to a "question-relay" mechanism: question states carry interfering cues alongside required-source evidence.
- Cutting pathways from the interfering modality to question states recovers more correct-answer logit than cutting those to the generation position.
- SECRET (SourcE-Conditioned RElay sTeering) is training-free and uses contrasting question representations from different modality-pathway interventions to steer question states toward required-source evidence.
- On CMM and AVHBench across three AVLLMs, SECRET beats prior training-free methods, with gains of up to +18.0 and +7.1 percentage points over base models.
Why it matters
Audio-visual models are asked to answer questions about one modality while another is also present, and the paper says cues from the unused modality can lead them to answers the required modality does not support. The authors argue that existing fixes work but that the internal cause was poorly understood. This work offers a specific mechanism, the question relay, and ties a method directly to it. Cutting pathways from the interfering modality to question states recovers more correct-answer logit than cutting those to the generation position, which is the evidence that the question states are where the interference builds up.
Who it affects
Mainly researchers and engineers working on audio-visual large language models and on hallucination mitigation, particularly those using training-free methods. The authors frame the underlying failure as a reliability problem for real-world applications of AVLLMs, so teams building on such models are the downstream audience.
How to use it
SECRET is described as training-free, so it works on the question states of an existing model rather than through retraining. It builds contrasting question representations through different modality-pathway interventions and uses them to steer the original question states toward required-source evidence. The authors also report that modality-specific captioning shows it carries over to open-ended generation. No code release, publication venue or submission date is mentioned in the abstract.
How solid is it
The evidence is the authors' own abstract. They report tests on two widely adopted benchmarks, CMM and AVHBench, across three AVLLMs, and say SECRET consistently outperforms prior training-free methods. The headline gains of up to +18.0 and +7.1 percentage points over base models are maximum figures, not averages, and the abstract does not say which benchmark or metric each belongs to or give baseline scores. The names of the three models are not given. The abstract names no authors or institutions.
Risks and caveats
The comparison is against prior training-free methods and base models; no comparison to trained (fine-tuned) methods is stated. No inference-time cost or latency of SECRET is stated. Because the reported gains are "up to" figures, typical improvement may be smaller. The mechanism and method are described only at abstract level here, so details of how the interventions are applied are not available.
“SECRET consistently outperforms prior training-free methods, substantially mitigating source-confused grounding hallucinations”
— Paper abstract