LMBuild benchmark tests whether LLM agents can design buildable structures

LMBuild is a benchmark for evaluating LLM agents on generating buildable and functional structures. The paper starts from a gap: LLM-based agents are increasingly capable of generating complex 3D structures, but producing elegant geometry is fundamentally different from producing objects that can be built and do their intended job. Existing evaluations, the authors say, largely focus on geometric quality while overlooking physical realizability.
LMBuild represents each generated object as an assembled structure made of part decompositions, joints, materials, and sequences. To make evaluation reproducible, the authors provide a unified framework with three components. The first is an interactive environment in which agents use tools to retrieve, create, and place components to construct objects. The second is a curated benchmark that repurposes established CAD datasets and augments them with knowledge from Wikipedia. The third is an evaluation framework covering four dimensions: structural soundness, functional affordance, design quality, and physical realization.
The authors report evaluations across 30 systems and describe three findings. (a) Soundness and alignment are no longer the primary bottlenecks for frontier closed-source models, while functional affordance and physical operability remain substantially more challenging. (b) Stronger models more effectively create new components, whereas weaker models tend to rely on retrieval. (c) Providing functional specifications substantially improves part completeness, kinematics, and physical operability.
From these results the authors conclude that generating real-world structures requires deeper reasoning about functional affordances, mechanics, and designing and creating novel components. They expect LMBuild to provide a foundation for measuring progress and incentivizing research toward agents that generate buildable and functional structures.
Key facts
- LMBuild is a benchmark that judges LLM agents on whether the structures they generate can be built and perform their intended function, not just on geometric quality.
- Objects are represented as assembled structures: part decompositions, joints, materials, and sequences.
- The framework has three parts: an interactive tool-using environment, a benchmark built from established CAD datasets plus Wikipedia knowledge, and an evaluation across structural soundness, functional affordance, design quality, and physical realization.
- Across 30 systems, soundness and alignment are no longer the main bottlenecks for frontier closed-source models, but functional affordance and physical operability remain much harder.
- Stronger models create new components more effectively, weaker ones lean on retrieval, and giving functional specifications substantially improves part completeness, kinematics, and physical operability.
Why it matters
Agents that output 3D structures are usually judged on how good the geometry looks. The LMBuild authors argue that this overlooks physical realizability: an elegant shape is not the same as an object that can be assembled and does its job. By scoring part decompositions, joints, materials, and assembly sequences, the benchmark moves evaluation closer to what it takes to make something real. Its headline finding is that frontier closed-source models have largely cleared soundness and alignment, while functional affordance and physical operability are still substantially harder.
Who it affects
The paper speaks to researchers building or evaluating LLM agents that design objects, and to anyone working with CAD-style generation who wants a measure beyond geometric quality. The framework also gives agent builders a way to compare how their systems handle retrieving, creating, and placing components. The authors say they hope it will incentivize research toward agents that generate buildable and functional structures.
How to use it
LMBuild is offered as a unified framework for reproducible evaluation: an interactive environment where an agent uses tools to retrieve, create, and place components, a curated benchmark drawn from established CAD datasets with Wikipedia knowledge added, and an evaluation framework covering four dimensions. One practical takeaway from the findings is that providing functional specifications substantially improves part completeness, kinematics, and physical operability. The text does not state a release date, code or dataset link, so check the paper page before planning to run it.
How solid is it
This is a benchmark paper, and the findings are the authors' own reading of their evaluation of 30 systems. The text gives no model names, scores, or percentages, so the size of the gaps cannot be judged from it. Which systems ranked where is not stated either. The claims are stated qualitatively: soundness and alignment are 'no longer the primary bottlenecks' for frontier closed-source models, and stronger models create components more effectively than weaker ones.
Risks and caveats
No scores or numeric results are given for any model or dimension, and the names of the CAD datasets and the size of the benchmark are not stated. It is also not stated whether structures were actually physically built or only evaluated in simulation, so 'physical realization' should be read as the name of an evaluation dimension, not as proof of real-world builds. The conclusions about frontier closed-source models come from the authors' framing and have not been independently reproduced in the text.
“Existing evaluations largely focus on geometric quality while overlooking physical realizability.”
— LMBuild paper abstract