VTR-Bench tests how well video generators render text, best model scores 0.250 WER

Video generation models can now produce clips whose visual quality approaches cinematic standards, but the authors of a new paper argue that one thing is routinely overlooked: text. Signs, ads and labels carry information in everyday scenes, and a generated video can look compelling, with lifelike subjects, while still rendering the words in the scene incorrectly. Existing benchmarks, the authors say, mostly judge visual quality, aesthetic appeal and physical plausibility, and pay limited attention to text.\n\nTo fill that gap they introduce VTR-Bench, a systematic benchmark for Visual Text Rendering in video generation. It places text inside concrete application scenarios, such as advertisements and scientific videos, using 300 carefully constructed prompts that span five scenario categories. The paper does not list all five categories in the text available here.\n\nThe evaluation is automated and aligned with human judgments. It scores two things separately: text fidelity, through carrier-specific transcription, and scene and motion requirements, through a prompt-specific chain of query.\n\nBeyond measurement, the authors propose a Keyframe-Guided Agentic Framework. In it, a Director agent coordinates image generation, video generation and visual evaluation, and uses visual feedback to steer iterative refinement and candidate selection.\n\nThey ran experiments on 11 state-of-the-art video generation models. The results show widespread difficulties in rendering scene text accurately. The best-performing model recorded an overall word error rate (WER) of 0.250, where lower is better. The authors also analyse the text rendering failures to characterise the challenges current models face. They conclude that visual text rendering is a key challenge for video generation and that their work shows a practical path toward improvement. Code is available at github.com/hardenyu21/VTR-Bench.
Key facts
- VTR-Bench is a benchmark for Visual Text Rendering in video generation, with 300 prompts across five scenario categories, including advertisements and scientific videos.
- Evaluation is automated with human alignments: text fidelity is scored via carrier-specific transcription, scene and motion requirements via a prompt-specific chain of query.
- Across 11 state-of-the-art models, the best one reached an overall word error rate of 0.250 (lower is better), and difficulties with scene text were widespread.
- The authors also propose a Keyframe-Guided Agentic Framework, where a Director agent coordinates image and video generation with visual evaluation to refine and select candidates.
- Code is public at github.com/hardenyu21/VTR-Bench.
Why it matters
Most video benchmarks reward how good a clip looks and how plausible its physics are. The authors point out that this misses text, which is how scenes convey information, so a video can score well and still show garbled words. VTR-Bench makes that failure measurable. Its finding that the best of 11 models reached a WER of 0.250 suggests the problem is far from solved.
Who it affects
Developers of video generation models get a test aimed at a specific weakness. The scenarios named by the authors, advertisements and scientific videos, point to people who need readable text in generated clips. Researchers working on evaluation can reuse the prompts and the pipeline.
How to use it
The code is available at https://github.com/hardenyu21/VTR-Bench. The benchmark supplies 300 prompts and an automated pipeline that scores text fidelity separately from scene and motion requirements. The paper also describes a Director-agent framework that iterates on keyframes and picks among candidates using visual feedback.
How solid is it
The numbers come from the authors' own experiments on 11 models, and the pipeline is described as aligned with human judgments. The text available here names no models, gives WER for the best model only, and reports no figure for how much the agentic framework improves results, so the claim of a practical path toward improvement rests on the authors' statement.
Risks and caveats
The 300-prompt set is modest in size, and the best-model WER of 0.250 is a single overall number without a per-model breakdown in the text available here. Automated transcription-based scoring depends on the pipeline's alignment with human judgment, which the authors report but this summary cannot verify. The agentic framework's gains are not quantified in the abstract.