RenderRank reranks documents as images, cutting input tokens 16.5-35.5%

RenderRank reranks documents as images, cutting input tokens 16.5-35.5%

A paper on Hugging Face introduces RenderRank, a reranker that judges how relevant a document is to a query from a picture of the document's text rather than from a sequence of text tokens. The idea starts from a simple observation: if document text is rendered as an image, a vision-language model can encode it as visual tokens, and that can shorten the input compared with feeding in the text itself.

The authors argue the saving matters most for reranking. In that task each query has to score several candidate documents, so every token saved is saved again on each candidate evaluation. RenderRank learns query-dependent relevance scoring from these compressed visual document representations, in place of the text token sequences that conventional text-based rerankers use.

Training has two stages. First, the relevance scores produced from visual inputs are aligned with those of a text-based teacher. Then the model is refined on the relative scores of positive and negative documents for the same query.

The paper reports two sets of results. Across 11 datasets from BEIR, RenderRank uses 16.5-35.5% fewer input tokens and reaches an average NDCG@10 of 55.96. That outperforms all evaluated text-based baselines below 4B parameters, and some larger models too.

On four long-document datasets, RenderRank reaches an average NDCG@10 of 88.27 with approximately half the average input token count of the evaluated text-based rerankers. In this setting it delivers 1.70x the highest average throughput of the evaluated baselines.

The authors conclude that compressed visual representations can support accurate document relevance scoring, offering an alternative to text token representations for reranking.

Key facts

  • RenderRank is a reranker that scores query-document relevance from document text rendered as images, encoded by a vision-language model as compressed visual tokens.
  • Across 11 BEIR datasets it uses 16.5-35.5% fewer input tokens and averages 55.96 NDCG@10, ahead of all evaluated text-based baselines below 4B parameters and some larger models.
  • On four long-document datasets it averages 88.27 NDCG@10 with approximately half the input tokens of the evaluated text-based rerankers.
  • In the long-document setting its throughput is 1.70x the highest average throughput of the evaluated baselines.
  • Training runs in two stages: align scores with a text-based teacher, then refine the relative scores of positive and negative documents for the same query.

Why it matters

Reranking scores many candidate documents for every query, so input length translates directly into cost and speed. RenderRank attacks that cost from an unusual angle: it turns document text into images and lets a vision-language model read them as a shorter run of visual tokens. The reported result is fewer tokens without giving up accuracy, and on long documents roughly half the input of the evaluated text-based rerankers along with 1.70x the top baseline throughput. The authors present it as an alternative to text token representations for reranking.

Who it affects

The work is most relevant to people building retrieval and search pipelines that rerank candidate documents, especially over long documents where input length dominates. It also speaks to researchers working on vision-language models and on compressing long inputs into visual tokens.

How to use it

The source is a paper abstract, and it does not describe a product or a ready-to-run tool. No code, model or dataset release is mentioned. The practical takeaway is the recipe: render documents as images, encode them with a vision-language model, first align scores with a text-based teacher, then refine on positive and negative documents for the same query.

How solid is it

The claims are the authors' own, drawn from the paper's abstract. The evaluation covers 11 BEIR datasets and four long-document datasets, with NDCG@10 as the accuracy measure. The abstract does not name the individual baselines or their sizes, and gives no baseline NDCG@10 figures to set against 55.96 or 88.27. It also does not say whether RenderRank beats text-based rerankers of 4B parameters and above on average, only that it beats some larger models. No hardware or setup for the throughput measurement is stated.

Risks and caveats

The throughput figure is relative to the highest average among the evaluated baselines in the long-document setting only, so it should not be read as a general speedup. The 16.5-35.5% token saving is reported for the 11 BEIR datasets, and the roughly half figure for the four long-document datasets. The accuracy lead is stated against evaluated baselines below 4B parameters; for larger models the claim is limited to some of them. The names of the four long-document datasets are not given.