Multi-vector visual document indices can be inverted back into pages, paper finds

Multi-vector visual document indices can be inverted back into pages, paper finds

A paper listed on Hugging Face Papers (2610.09920) argues that the vector index behind multi-vector visual document retrieval is far more revealing than it looks. Prevailing retrievers of this kind store each page as about a thousand patch vectors, often in vector databases run by a third party. Because no one can read a page from its vectors, the index is easily treated as less sensitive than the page itself.

The authors hypothesize otherwise. The index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents. So whoever runs or breaches the store, they hypothesize, can reproduce a page from its index alone. They frame inversion as conditional document image generation. From the vectors the attack infers what it needs: the encoder, the page shape and, for shuffled vectors, their order.

On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, the inverted pages rank their source page first 98.4% of the time.

The authors then test two cheap protections: token pooling and shuffling. Both cut word recall to about 8%. Shuffling does not hold up, though. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%. Inverting a pooled index, by contrast, remains an open problem.

To test generalisation, the authors apply the same attack unchanged to another multi-vector retriever. Its inverted pages still rank their source page first 70.2% of the time, although its word recall stays below a nearest-neighbour baseline. Their conclusion is that multi-vector visual document retrievers are vulnerable to inversion through their stored index, which should be protected like the documents it encodes.

Key facts

  • On ViDoRe v3, pages inverted from raw multi-vector indices recover 47% of the words and 45% of the sensitive tokens.
  • Used as queries against the stored indices, the inverted pages rank their source page first 98.4% of the time.
  • Token pooling and shuffling both cut word recall to about 8%, but a model that restores the order of a shuffled index lifts the first-rank share from 3.8% to 93.5%.
  • Inverting a pooled index is described as still open; the same attack on another multi-vector retriever ranks the source page first 70.2% of the time, with word recall below a nearest-neighbour baseline.
  • The authors say the stored index should be protected like the documents it encodes.

Why it matters

Multi-vector visual document retrievers keep about a thousand patch vectors per page, and the common assumption is that a page cannot be read from its vectors. The paper challenges that assumption with numbers: nearly half the words and 45% of the sensitive tokens come back from a raw index on ViDoRe v3. The index is often hosted by a third party, so a copy of it may sit outside the organisation that owns the documents.

Who it affects

Anyone who builds or runs retrieval over page images with a multi-vector visual document retriever, especially where the vector database is run by a third party. The authors name whoever runs or breaches the store as the party able to reproduce a page. Owners of confidential documents indexed this way are the ones whose data is exposed.

How to use it

This is a security finding rather than a tool. The practical takeaway the authors give is to protect the stored index like the documents it encodes, not to treat it as a harmless derivative. No code, model release or dataset release is mentioned.

How solid is it

The figures come from the paper's abstract and a benchmark experiment on ViDoRe v3, plus one test on a second retriever. The abstract names no authors or institutions, and the retrievers tested are not identified. The order-restoring result is the strongest evidence against shuffling: the first-rank share goes from 3.8% to 93.5%. The conclusion that retrievers are vulnerable is the authors' own, drawn from their experiments.

Risks and caveats

Inversion is partial: 47% word recall means more than half the words are not recovered. Token pooling and shuffling cut word recall to about 8%, and inverting a pooled index remains open, so pooling may be a real obstacle. Shuffling is weak, since order can be restored. On the second retriever the same attack ranked the source page first 70.2% of the time, but its word recall stayed below a nearest-neighbour baseline. No real-world breach or attack is reported, and no effective defence is offered.

“Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.”

— Paper abstract, Hugging Face Papers 2610.09920