AnswerMap shows where a VLM looks using only its yes/no answers

AnswerMap shows where a VLM looks using only its yes/no answers

A paper on Hugging Face introduces AnswerMap, a method for showing which parts of an image a vision-language model (VLM) relies on when it answers a visual query. The authors start from a complaint about current tools. Text rationales use a mismatched modality, and internal read-outs originate too early to reflect the final output and require white-box access to the model.

AnswerMap is described as a training-free, task-agnostic, black-box visual rationale constructed from the output head. The image is cut into K row bands and K column bands. Each band is shown alone to the frozen model along with the query, phrased as a yes/no relevance question. The outer product of the row and column "yes" posteriors gives the query-conditioned spatial map.

The authors then define a fixed read-out R on top of the map, for example an expectation or a maximum. That lets them derive continuous outputs such as a location natively. They say this bypasses reliance on discrete text tokens for continuous-output tasks and guarantees an image-dependent answer by construction.

Because a rationale can be confabulated, the authors validate the map across four models and three query distributions with two tests: agreement with the model's own generated point, and deletion of the map's region. The map lands where the model points, with AUC 0.85 against 0.38 for attention. Deleting its region flips 53% of correct answers, against 19% when attention's region is deleted.

Beyond faithfulness, the paper shows the map's task-agnostic use through three read-outs. Its maximum flags hallucinated objects without generation. Its expectation localizes correctly when the model's own pointing fails. Its top-mass region, fed back as a crop, fixes half of the model's wrong answers. The authors conclude that AnswerMap offers a new lens on VLM interpretability and, through its read-outs, a new output interface for visual tasks beyond text tokens.

Key facts

  • AnswerMap is a training-free, task-agnostic, black-box method: the image is cut into K row and K column bands, each shown alone to the frozen VLM as a yes/no relevance question, and the outer product of the "yes" posteriors forms the spatial map.
  • Agreement with the model's own generated point: AUC 0.85 for AnswerMap against 0.38 for attention.
  • Deleting the AnswerMap region flips 53% of correct answers, against 19% for attention's region.
  • Validation covers four models and three query distributions, using two tests: pointing agreement and region deletion.
  • Three read-outs are demonstrated: the maximum flags hallucinated objects without generation, the expectation localizes when the model's own pointing fails, and the top-mass region fed back as a crop fixes half of the model's wrong answers.

Why it matters

The authors argue that existing ways of explaining a VLM answer fall short. Text rationales sit in a different modality from the image, and internal read-outs come too early in the network to reflect the final output while needing white-box access. AnswerMap instead builds a visual rationale from the output head alone, so it works from the model's own yes/no answers. The paper also treats the map as more than an explanation: read-outs on top of it give continuous outputs such as location without going through discrete text tokens, which the authors present as a new output interface for visual tasks.

Who it affects

The method is aimed at people who study or debug VLMs and want to see which image regions drive an answer. Because it is black-box, it needs no white-box access to the model, which the authors name as a limit of internal read-outs. The three demonstrated read-outs point to practical uses: flagging hallucinated objects, localizing when the model's own pointing fails, and correcting wrong answers by cropping.

How to use it

The recipe in the abstract is simple. Cut the image into K row bands and K column bands. Show each band alone to the frozen model with the query, asked as a yes/no relevance question. Take the outer product of the row and column "yes" posteriors to get the spatial map. Then apply a fixed read-out R: the maximum to flag hallucinated objects, the expectation to get a location, or the top-mass region fed back as a crop to improve answers. No code release, dataset or venue is mentioned, and the value of K is not given.

How solid is it

The authors address faithfulness directly, noting that a rationale can be confabulated. They test AnswerMap across four models and three query distributions. The map agrees with the model's own generated point at AUC 0.85 against 0.38 for attention, and deleting its region flips 53% of correct answers against 19% for attention's region. These are the authors' own reported results from an abstract. The four models and three query distributions are not named.

Risks and caveats

It is not stated whether the 53% and 19% deletion figures are averages over models or per model, nor whether "fixes half of the model's wrong answers" is an average across models or a specific model. No compute cost or number of forward passes per image is stated. The value of K is not given. The models and query distributions behind the results are not named, so it is hard to judge how far the numbers generalize.

“The map lands where the model points (AUC 0.85 against 0.38 for attention), and deleting its region flips 53% of correct answers (against 19% for attention's).”

— AnswerMap paper abstract