Last Translation Benchmark targets machine translation models that pass every existing test

A paper by Vilém Zouhar and co-authors argues that machine translation evaluation has stalled: standard benchmarks are approaching saturation as models get stronger, automatic translation metrics are unreliable and vulnerable to reward hacking, and even human evaluation, long treated as the gold standard, is not problem free because it often lacks reproducibility, objectivity and scalability. Together these gaps make it hard to track real progress in the field or see where models still fail. The authors respond with the Last Translation Benchmark (LTB), a collection of human-authored, peer-reviewed examples spanning text, images, audio and video, each one chosen specifically because it breaks leading machine translation models. Alongside the dataset, they introduce a new evaluation approach: every example ships with handcrafted verification rules that describe the concrete failure case it targets, so future evaluation can check a model's output against a specific, actionable rule rather than relying on a fuzzy automatic score or an unreproducible human judgment. LTB is designed as a live dataset that keeps accepting new contributions rather than shipping once and going stale. The current release, LTBv1, contains every contribution accepted before September 1st 2026, and the authors say further releases are planned as more data is continuously collected. The paper does not name the authors' institutions, does not state how many examples LTBv1 contains, reports no benchmark scores for any specific model, and does not detail who authors or peer-reviews new contributions or how that process works.
Key facts
- The Last Translation Benchmark (LTB) is built from human-authored, peer-reviewed examples across text, images, audio and video, each selected because it breaks leading machine translation models.
- Each example carries handcrafted verification rules describing its specific failure case, aimed at making evaluation reliable and actionable rather than a single fuzzy score.
- The authors motivate the work by pointing to three separate problems: standard MT benchmarks nearing saturation, automatic metrics that are unreliable and reward-hackable, and human evaluation that often lacks reproducibility, objectivity and scalability.
- LTB is a live dataset that keeps accepting contributions; the current release, LTBv1, covers everything accepted before September 1st 2026, with further releases planned.
- The paper does not report how many examples LTBv1 contains, does not give benchmark results for any model, and does not detail the peer-review process for new contributions.
Why it matters
Machine translation evaluation has a saturation problem: as models improve, the benchmarks meant to test them stop discriminating between a strong model and a great one, and the automatic metrics used at scale are both unreliable and gameable. Human evaluation was supposed to be the fallback, but the authors point out it is often not reproducible, not objective and does not scale. LTB is an attempt to fix the measurement problem itself, not just add another leaderboard: it targets specifically the inputs that current top models still fail on, and it attaches a checkable rule to each one instead of a single aggregate score.
Who it affects
Researchers and engineers building or evaluating machine translation systems, and teams building evaluation pipelines for multilingual or multimodal AI more broadly, since the dataset spans text, images, audio and video rather than text alone.
How to use it
The paper describes LTB as a live dataset open to ongoing contributions, with LTBv1 the current release covering material accepted before September 1st 2026 and further releases planned as more data comes in. The source text gives no pricing, license terms or access instructions, and does not state where the dataset or its verification rules can be obtained.
How solid is it
The paper is presented as peer-reviewed at the level of individual examples: each item in the dataset is described as human-authored and peer-reviewed before inclusion, and each carries a handcrafted verification rule rather than relying on an automatic metric. The source text, however, gives no benchmark scores for any model, no count of how many examples LTBv1 contains, and no detail on the composition of contributors or reviewers, which limits how much can be judged from the abstract alone.
Risks and caveats
The abstract names no authors' institutions, gives no size for the dataset, reports no results demonstrating how leading models actually perform against LTB, and does not explain who authors or peer-reviews new contributions or by what process. A benchmark's value depends heavily on the rigor of that process and on how it is maintained over time as a live, continuously growing dataset; none of that is verifiable from the material available here.