Automation & Agents · 3 Oct 2026 · 10:31 CEST
AI agents build 3D scenes from photos but have no idea if they got it right

Publisher preview · OZZZER analysis pending editorial review.
Coding agents can build an editable 3D scene from a single photo by writing and refining code step by step. The benchmark that comes with this work reveals where the process still breaks down: self-assessment and geometric accuracy. A 3D reconstruction from a single photo is most useful when it exists as an executable program you can inspect, edit, and query.
That's the premise behind LEGO-Anything, a project from researchers at the University of Maryland and AWS. The approach is called "Image-to-Code." A coding agent receives a single image and writes code for Blender, the widely used 3D software. Rather than generating the scene in one pass, the agent works iteratively: it writes code, runs it, looks at the result, and revises until the scene matches the original.
Because the output is a program, it captures objects, geometry, layout, and camera position explicitly. You can run, check, and modify the scene like any other piece of code. To measure how well agents perform, the team introduces LEGO-Bench. It contains 208 images from 104 indoor and outdoor scenes and uses 443…
Excerpt supplied by the publisher.
Source
THE DECODER · 3 Oct 2026 · 10:31 CEST
Open the original at THE DECODER ↗