Researchers at the University of Maryland and AWS have developed LEGO-Anything, a system in which a coding agent reconstructs a photograph as an editable Blender program. The agent writes code, renders the scene, inspects the result and revises objects, geometry, layout and camera position.

The accompanying LEGO-Bench uses 208 images from 104 simulator scenes, giving researchers exact hidden geometry rather than an uncertain estimate from real photographs. All six tested GPT configurations usually produced a valid artifact, but faithful reconstruction was much harder.

GPT-6 Astra led the test with 53.4 percent accuracy on indoor scenes and 39.6 percent outdoors. More reasoning improved some results, yet agents often damaged a good partial scene during later revisions. When asked which of two versions had better geometry, their choices were around chance level.

A training-free plugin improved performance by anchoring the initial scene, using concrete measurements instead of model self-judgment and protecting correct work from regressive edits. Weaker agents gained the most, in some cases by as much as 62.7 percent. The result suggests executable 3D reconstruction is useful for inspection and editing, but current agents still need external measurements before their scenes can be trusted for geometric analysis.