EngiWorld benchmark: best of seven frontier models scores only 44.3 on engineering software tasks

The authors of EngiWorld argue that autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach. Their explanation is that engineering workflows demand reasoning over geometric and physical constraints, and over dependencies that must be preserved across software and design stages.
To measure this, they present EngiWorld, which they call the first benchmark structured around the complete design loop. It contains 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms. Agents can work through both GUI and CLI interfaces, and the tasks come in 6 types, ranging from software-selection to open-ended tasks.
The authors also introduce an artifact-centric evaluation methodology. It rests on a unified domain-verifier suite that programmatically checks the geometric validity, physical feasibility and rule compliance of both final and intermediate artifacts. Quantitative design tasks are scored continuously by specification attainment rather than as a binary pass or fail.
The authors evaluated seven frontier models and report a substantial capability gap. The strongest model reaches an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. They present EngiWorld as the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.
Key facts
- EngiWorld has 1,301 expert-curated tasks across 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms.
- Tasks use both GUI and CLI interfaces and come in 6 types, from software-selection to open-ended tasks.
- A domain-verifier suite programmatically checks geometric validity, physical feasibility and rule compliance of final and intermediate artifacts.
- Among seven evaluated frontier models, the strongest achieves an EngiScore of only 44.3.
- Just 3.6% of multi-software attempts succeed.
Why it matters
General-purpose computer-use agents are improving quickly, but the authors say professional industrial engineering is a harder test: work depends on geometric and physical constraints, and on dependencies that carry across software and design stages. They describe EngiWorld as the first benchmark structured around the complete design loop. The headline result is a gap: the best of seven frontier models scores 44.3 on EngiScore, and only 3.6% of multi-software attempts succeed.
Who it affects
The benchmark targets researchers and developers building agents that operate professional engineering software. Its six domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) point to engineers and teams working in mechanical design, simulation, manufacturing, building information modelling, electronic design and 3D visualization, who would be the eventual users of such agents.
How to use it
The authors position EngiWorld as a foundation for measuring progress toward agents that operate professional engineering software end to end. The abstract does not state whether the benchmark or code is publicly released. Practitioners should treat it for now as a yardstick for agent capability rather than a tool they can run.
How solid is it
The claims come from the authors' own abstract. The evaluation design is concrete: 1,301 expert-curated tasks, programmatic verifiers for intermediate and final artifacts, and continuous scoring of quantitative tasks by specification attainment. The 44.3 figure is the maximum among seven models, not an average. The abstract does not name the seven frontier models or which one scored 44.3, and it does not give the scale or maximum of EngiScore. It also gives no per-domain, per-interface (GUI vs CLI) or per-task-type results.
Risks and caveats
The abstract does not say how many multi-software attempts were made, nor the success rate of single-software attempts, so the 3.6% figure is hard to put in context. It does not name the 26 software platforms or the authors and institutions. Because the benchmark is new and the claim of being the first is the authors' own, the headline numbers should be read as one team's measurement until others reproduce them.
“reliable automation of professional industrial engineering remains out of reach”
— EngiWorld abstract