BrickBench: Evaluating Agentic Brick Design
We evaluate coding agents on their ability to design LEGO assemblies from a text prompt across three settings with different build constraints: Model (at most 400 parts), Set (400–4000 parts), and Alt-Build (the part inventory of set 10698). Each of the assemblies shown below is composed of actual LEGO parts and can be physically built. Click on a build to view it in the gallery and rotate it.
Task-specific ApproachesA general coding agent effectively saturates prior evaluations
BrickNet and BrickGPT finetune language models to autoregressively assemble brick structures. We evaluate GPT-5.6 Luna, a general coding agent given our environment, on the BrickNet validation set.
| Method | PE ↑ | SigLIP 2 ↑ | VQAScore ↑ |
|---|---|---|---|
| BrickGPT | 0.157 | 0.052 | 0.050 |
| BrickNet-0.6B | 0.279 | 0.603 | 0.557 |
| BrickNet-1.7B | 0.282 | 0.631 | 0.593 |
| BrickNet-4B | 0.283 | 0.639 | 0.615 |
| BrickNet-8B | 0.284 | 0.647 | 0.608 |
| BrickNet-14B | 0.283 | 0.625 | 0.602 |
| GT | 0.315 | 0.826 | 0.748 |
| GPT-5.6 Luna | 0.313 | 0.818 | 0.732 |
Sample inputs and outputs from BrickNet’s validation set.
Caption alignment following BrickNet’s evaluation protocol.
Without task-specific training, Luna outperforms both, within 0.02 of the held-out assemblies on each metric. BrickNet evaluates small objects of at most 100 parts, far from the complexity of a full LEGO set. This motivates our design of a new benchmark.
The benchmarkBrickBench: Evaluating Agentic Brick Design
Given a text prompt, an agent must produce an assembly that satisfies semantic and design criteria and that can be physically built. It adds two challenges prior evaluations leave unexplored, physical complexity and design quality. It consists of 300 prompts over three constraint settings:
Model
At most 400 parts. Assemblies of primarily a single component.Browse in the gallery →Set
400–4000 parts. An assembly on the scale of a full-size retail set.Browse in the gallery →Alt-Build
Only the 783 pieces of retail set 10698. A scarce, fixed part inventory.Browse in the gallery →We create ten prompts in each of ten categories, aligned with the theme classifications of the BrickLink Designer Program: Medieval / Castle, Vehicle / Boat / Airplane, Train, Space / Sci-Fi / Fantasy, Art / Object, Building, Floral / Nature, Animal, Pirate, and History / Period.
How we score it
- Valid: meets the part requirement, collision-free, and stable.
- VQA: the fraction of the prompt’s requirements satisfied, split into Yes/No questions and evaluated by a VLM.
- ELO: Bradley–Terry ratings from pairwise VLM judgments, of semantic alignment with the prompt (Align ELO) and of design agnostic of it (Design ELO).
- Cost: the list price of the tokens consumed per assembly in USD.
Question graph
The environmentBrickAgent: Agentic LEGO Design
BrickAgent is the environment of BrickBench. It gives agents tools to programmatically construct, inspect, and validate LEGO assemblies, placing parts through the connector system of BrickNet.
ResultsPrompt adherence saturates; headroom lies in design quality
We evaluate eleven frontier agents and two data-driven baselines. Five agents deliver a valid assembly for every prompt. VQA separates them gently and Design ELO much more sharply.
| System | Valid | VQA | ELO | Align ELO | Design ELO | Cost ($) |
|---|---|---|---|---|---|---|
Metrics per setting and in the aggregate (Overall). Higher is better, except for Cost; ELO intervals are 95%. The two data-driven baselines apply only to the Model setting.Tap a column to sort.
(a) ELO against mean assembly cost.
(b) Design ELO against human judgments of design quality, for the nine reference agents.
(a) Rank does not follow price: GPT-6.1 Sol matches Astra on ELO at a fifth of its cost. (b) Design ELO agrees with human raters on 34 of 36 agent pairs (Kendall τ = 0.89).
Environment ablationRemoving the environment sharply reduces physical validity
Share of valid assemblies over the 300 prompts, with BrickAgent or with only the LDraw library and a primer on format.
| System | Valid ↑ | Collisions ↓ | Components | VQA ↑ | ELO ↑ |
|---|---|---|---|---|---|
| 1.00 | 0.00 | 2.21 | 0.954 | 1297 ±22 | |
| without BrickAgent | 0.40 | 7.19 | 6.12 | 0.940 | 1309 ±24 |
| 1.00 | 0.00 | 4.60 | 0.792 | 898 ±16 | |
| without BrickAgent | <0.01 | 155.51 | 83.00 | 0.813 | 960 ±20 |
Pooled over the three settings (300 prompts). Components are connected components, each per assembly.
Astra remains valid on 40% of prompts and Luna on one in 300. However, alignment and design do not fall: without the restriction of buildability, Luna makes assemblies that depict the prompt as well and look better, but that cannot be built.
DesignEqually valid, yet differing dramatically in design
The agents vary in their interpretation of the prompts. ClickTap any build to open it in the gallery.
The ceilingA floor a simulator can verify, an open ceiling on design quality
Physical validity
Connected? Collision-free? Stable under gravity?Design quality
Clever part usage, shapes, proportions, and details.We asked people familiar with LEGO design to tell agent builds from human-designed models of a similar part count, as a sort of "LEGO Turing test."
How often raters took each system’s build for the human-designed one; 0.5 is chance.
Human-designed models from the BrickNet dataset.
Raters consistently found the human build. No agent was taken for human more than one time in five.
What remainsAgents learn to build what works, but not yet what makes a great design
Design quality
Part usage, proportion, and surface detail.Physical validity and prompt adherence
Connections, collisions, stability, and the prompt.Paper
Abstract
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch.
BibTeX
@misc{kulits2026brickbenchevaluatingagenticbrick,
title={BrickBench: Evaluating Agentic Brick Design},
author={Peter Kulits and Yiqing Xu and R. Kenny Jones and Cordelia Schmid and Jiajun Wu},
year={2026},
eprint={2610.12452},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2610.12452},
}














