ColNanoVDR shrinks multi-vector document retrieval query encoders without using any pages

Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. The paper's authors argue that distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck.
The standard recipe for that distillation matches the teacher's MaxSim scores. That means every training page has to be encoded and cached, which can reach terabytes of page tokens. An earlier method, NanoVDR, avoids pages entirely by training on the teacher's query embeddings alone, but it works only for single-vector retrievers.
The paper presents ColNanoVDR, which the authors call, to their knowledge, the first framework to bring this document-free distillation to multi-vector VDR. Its objective is OTW (Optimal Transport with Learned Weights). OTW aligns the student's query tokens with the teacher's by entropic optimal transport, with a learned weight for each student token, and it needs no correspondence between the two tokenizations. The authors also prove that the resulting alignment cost bounds the MaxSim score difference on every page.
The reported results: distilled from five state-of-the-art teachers, the 149M text-only students retain about 95% of their teachers' NDCG@5 on ViDoRe v1-v3 while encoding queries up to 26x faster. Under identical training, OTW matches score distillation while encoding no page and reading 12.6x less cached teacher data.
Key facts
- ColNanoVDR distills a multi-vector VDR query encoder into a small text-only student using only the teacher's query embeddings, with no document pages.
- The OTW objective (Optimal Transport with Learned Weights) aligns student and teacher query tokens by entropic optimal transport, with no need for matching tokenizations.
- The authors prove the alignment cost bounds the MaxSim score difference on every page.
- Students of 149M parameters, distilled from five teachers, retain about 95% of teacher NDCG@5 on ViDoRe v1-v3 and encode queries up to 26x faster.
- Under identical training, OTW matches score distillation while reading 12.6x less cached teacher data.
Why it matters
Multi-vector retrievers lead visual document retrieval, yet each search runs a multi-billion-parameter query encoder. A small student that works against the teacher's existing index would remove that cost. The usual way to train one, matching MaxSim scores, needs every training page encoded and cached, which can reach terabytes of page tokens. The authors say ColNanoVDR is, to their knowledge, the first framework to make document-free distillation work for multi-vector VDR; the earlier NanoVDR did this only for single-vector retrievers.
Who it affects
The work is aimed at people who build or run multi-vector visual document retrieval systems on vision-language models and want cheaper query encoding. It also matters to anyone who wants to distill such a retriever but cannot afford to encode and cache a large page corpus for training.
How to use it
The source describes the method, not a release. It mentions no code or model release, license or availability. In practice the recipe is this: take a multi-vector teacher, collect its query embeddings, and train a small text-only student with the OTW objective. The student then queries the teacher's existing index, so no pages are encoded during training.
How solid is it
The source is a paper abstract, so the results are the authors' own. Two claims stand out: a proof that the alignment cost bounds the MaxSim score difference on every page, and results from five state-of-the-art teachers on ViDoRe v1-v3. The 95% figure is approximate ('about') and is relative to teacher NDCG@5; the source gives no absolute scores. It also does not break the figure down per teacher or per ViDoRe version. The speedup is stated as 'up to 26x', so it is a maximum. The five teachers are not named in the source.
Risks and caveats
The students give up roughly 5% of their teachers' NDCG@5, on the authors' figures. The 26x speedup is a ceiling, and the source does not say whether it is typical. The claim of being the first document-free distillation framework for multi-vector VDR is the authors' own, made to their knowledge. The source names no authors or institutions and does not say whether the work is independently reproduced.
“to our knowledge the first framework to bring this document-free distillation to multi-vector VDR”
— ColNanoVDR paper abstract