The strongest of seven AI agents completed only 17.5% of tasks in CADWorld, a new benchmark for operating professional computer-aided design software. Human experts passed 87% of the same work, exposing a large gap between general desktop interaction and reliable engineering output.
CADWorld contains 200 long-horizon tasks across 11 FreeCAD workflow categories, including sketching, part modeling, assemblies, computer-aided manufacturing, finite-element analysis, measurement, mesh processing and technical drawing. Agents see screenshots and act through the graphical interface. Evaluation then runs task-specific checks against the saved native project and supporting files, testing dimensions, geometry, constraints, parametric structure, manufacturing state and simulation results. A visually plausible screenshot is not enough.
Weaker systems frequently failed before creating a valid artifact. Stronger agents got further but still violated structural or construction-process requirements that affect whether an engineer can safely edit or manufacture the result. The benchmark covers one open-source CAD package and does not establish performance in every engineering tool. Its key contribution is durable verification: agents must leave behind a correct, structured project rather than merely appear to complete a sequence of clicks. That standard is essential for professional automation where hidden model structure matters as much as the rendered shape.