EditHero benchmark: LLM agents keep unedited 3D parts intact but take minutes per edit

EditHero benchmark: LLM agents keep unedited 3D parts intact but take minutes per edit

3D editing methods are usually tested on a single edit, yet a real asset is built through a long sequence of revisions. Each revision has to implement the requested change and leave everything else unchanged. EditHero is a new benchmark aimed at that gap. Its authors describe it, to their knowledge, as the first benchmark for long-horizon, part-level 3D editing. It comes with natural-language instructions and target images for both geometry and texture.\n\nTo keep the ground truth exact, a deterministic assembly engine produces the precise target after every edit, and every edit sequence is reviewed by hand.\n\nThe authors use EditHero to compare 2 opposite approaches. Non-agentic methods work top down: they regenerate the object from a learned 3D representation and infer what to keep. LLM/VLM agents work bottom up: they edit through code that inspects the mesh and rewrites only the parts the instructions require.\n\nThe results point in different directions. The non-agentic methods often miss the requested change and disturb regions that should stay fixed. Most LLMs follow instructions more closely, and all of them preserve the unedited parts better, but each of their edits takes minutes. The authors say they will release the engine and the edit sequences to support research on reliable iterative 3D editing.

Key facts

  • EditHero is described by its authors as, to their knowledge, the first benchmark for long-horizon, part-level 3D editing, with natural-language instructions and target images for geometry and texture.
  • A deterministic assembly engine produces the exact target after every edit, and every sequence is reviewed by hand.
  • It compares 2 approaches: non-agentic methods that regenerate the object top down, and LLM/VLM agents that edit bottom up through code that inspects the mesh.
  • Non-agentic methods often miss the requested change and disturb regions that should stay fixed.
  • Most LLMs follow instructions more closely and all preserve unedited parts better, but each edit takes minutes.

Why it matters

Existing 3D editing tests usually check a single edit, while real assets are built through many revisions, each of which must change one thing and leave the rest alone. EditHero tests that long-horizon, part-level setting directly. It also sets two philosophies side by side: regenerating an object from a learned 3D representation and inferring what to keep, versus writing code that rewrites only the required parts. The comparison shows a trade-off between preserving what should not change and speed.

Who it affects

Researchers building 3D editing methods get a benchmark that scores sequences of edits rather than one-off changes, with targets for both geometry and texture. Teams weighing code-writing LLM/VLM agents against non-agentic editors for 3D asset work get a comparison of how each handles instruction following and unedited regions.

How to use it

The authors say they will release the assembly engine and the edit sequences to support research on reliable iterative 3D editing. The source gives no release date or repository link for them. Once available, they could be used to test a 3D editing method on multi-step edit sequences against exact targets for geometry and texture.

How solid is it

The claims come from the authors' own abstract. The benchmark rests on a deterministic assembly engine that yields the exact target after every edit, and every sequence is reviewed by hand. The first-benchmark claim is hedged by the authors themselves as 'to our knowledge'. The source text gives no numeric scores, no benchmark size and no named models or methods, so the findings are qualitative.

Risks and caveats

The headline finding is a trade-off: LLM agents preserve unedited parts better, but each of their edits takes minutes, and the source gives no finer timing. Only most LLMs, not all, follow instructions more closely than the non-agentic methods. The source does not say which LLMs fell short. The comparison covers two broad approaches, and the source names no specific systems, so it is hard to say how the results apply to any particular tool.

“The non-agentic methods often miss the requested change and disturb regions that should stay fixed.”

— EditHero abstract