BrickBench: Evaluating Agentic Brick Design

Peter Kulits1,2 Yiqing Xu1 R. Kenny Jones1 Cordelia Schmid3 Jiajun Wu1
1Stanford University 2Max Planck Institute for Intelligent Systems 3Inria
arXiv Code PyPI Gallery

We evaluate coding agents on their ability to design LEGO assemblies from a text prompt across three settings with different build constraints: Model (at most 400 parts), Set (400–4000 parts), and Alt-Build (the part inventory of set 10698). Each of the assemblies shown below is composed of actual LEGO parts and can be physically built. Click on a build to view it in the gallery and rotate it.

SetA brown dog with floppy black ears and a small brown hat sits at a table in a burning room. A white coffee mug rests in front of it. Flames climb the curtains and surround the chair, but the dog sits upright with its paws resting calmly on the tabletop. The front of the room is open to reveal the scene.
Alt-BuildA cowboy in a wide-brimmed hat sits on a saddled horse beside a cactus with two raised arms. He holds a looped lasso above his head. The horse’s front legs stand on a low rock, while its rear legs remain on the sandy ground.

Task-specific ApproachesA general coding agent effectively saturates prior evaluations

BrickNet and BrickGPT finetune language models to autoregressively assemble brick structures. We evaluate GPT-5.6 Luna, a general coding agent given our environment, on the BrickNet validation set.

A small robot: the held-out assembly and three systems' outputs
This is a small, blocky LEGO robot model, primarily constructed from yellow and black bricks, featuring articulated limbs with wheel-like hands and feet, and presented from multiple angles to showcase its design.
A minifigure: the held-out assembly and three systems' outputs
This is a 3D model of a classic LEGO minifigure, featuring a yellow head and hands, a red baseball cap, a plain white torso, and blue legs, shown from multiple angles against a white background.
MethodPE ↑SigLIP 2 ↑VQAScore ↑
BrickGPT0.1570.0520.050
BrickNet-0.6B0.2790.6030.557
BrickNet-1.7B0.2820.6310.593
BrickNet-4B0.2830.6390.615
BrickNet-8B0.2840.6470.608
BrickNet-14B0.2830.6250.602
GT0.3150.8260.748
GPT-5.6 Luna0.3130.8180.732

Sample inputs and outputs from BrickNet’s validation set.

Caption alignment following BrickNet’s evaluation protocol.

Without task-specific training, Luna outperforms both, within 0.02 of the held-out assemblies on each metric. BrickNet evaluates small objects of at most 100 parts, far from the complexity of a full LEGO set. This motivates our design of a new benchmark.

The benchmarkBrickBench: Evaluating Agentic Brick Design

Given a text prompt, an agent must produce an assembly that satisfies semantic and design criteria and that can be physically built. It adds two challenges prior evaluations leave unexplored, physical complexity and design quality. It consists of 300 prompts over three constraint settings:

Model

At most 400 parts. Assemblies of primarily a single component.Browse in the gallery →

Set

400–4000 parts. An assembly on the scale of a full-size retail set.Browse in the gallery →

Alt-Build

Only the 783 pieces of retail set 10698. A scarce, fixed part inventory.Browse in the gallery →

We create ten prompts in each of ten categories, aligned with the theme classifications of the BrickLink Designer Program: Medieval / Castle, Vehicle / Boat / Airplane, Train, Space / Sci-Fi / Fantasy, Art / Object, Building, Floral / Nature, Animal, Pirate, and History / Period.

How we score it

Question graph
Is there a recognizable dog? Is the dog brown? Are the dog's ears floppy? Are the dog's ears black? Is the dog wearing a hat? Is the dog sitting at the table? Is the mug in front of the dog? Is the dog sitting upright? Are the dog's paws resting on the tabletop? Is the hat brown? Is there a room? Are there flames in the room? Is the table in the room? Does the room have curtains? Is one side of the room open so that the scene inside is visible? Do the flames climb the curtains? Do the flames surround the chair? Is there a table? Is there a coffee mug? Is there a chair?
The question graph for the calm dog burning prompt above. Questions are colored by type: entity, attribute, or relation. An arrow from one question to another means the second is satisfied only if the judge answers yes to the first. HoverTap a question to trace what it depends on and what depends on it.

The environmentBrickAgent: Agentic LEGO Design

BrickAgent is the environment of BrickBench. It gives agents tools to programmatically construct, inspect, and validate LEGO assemblies, placing parts through the connector system of BrickNet.

ResultsPrompt adherence saturates; headroom lies in design quality

We evaluate eleven frontier agents and two data-driven baselines. Five agents deliver a valid assembly for every prompt. VQA separates them gently and Design ELO much more sharply.

SystemValidVQAELOAlign ELODesign ELOCost ($)

Metrics per setting and in the aggregate (Overall). Higher is better, except for Cost; ELO intervals are 95%. The two data-driven baselines apply only to the Model setting.Tap a column to sort.

(a) ELO against mean assembly cost.

(b) Design ELO against human judgments of design quality, for the nine reference agents.

(a) Rank does not follow price: GPT-6.1 Sol matches Astra on ELO at a fifth of its cost. (b) Design ELO agrees with human raters on 34 of 36 agent pairs (Kendall τ = 0.89).

Environment ablationRemoving the environment sharply reduces physical validity

Share of valid assemblies over the 300 prompts, with BrickAgent or with only the LDraw library and a primer on format.

SystemValid ↑Collisions ↓ComponentsVQA ↑ELO ↑
GPT-6 Astra1.000.002.210.9541297 ±22
without BrickAgent0.407.196.120.9401309 ±24
GPT-5.6 Luna1.000.004.600.792898 ±16
without BrickAgent<0.01155.5183.000.813960 ±20

Pooled over the three settings (300 prompts). Components are connected components, each per assembly.

Astra remains valid on 40% of prompts and Luna on one in 300. However, alignment and design do not fall: without the restriction of buildability, Luna makes assemblies that depict the prompt as well and look better, but that cannot be built.

DesignEqually valid, yet differing dramatically in design

The agents vary in their interpretation of the prompts. ClickTap any build to open it in the gallery.

ModelA green dumpster on four small black wheels has its lid pushed open by a tall cluster of red, orange, and yellow flames. A crumpled white sheet protrudes from one corner, and a tipped-over trash can lies beside it.
ModelA cramped mailroom corner has a wall covered in white notices, envelopes, and photographs joined by tangled red strings. A wild-eyed, brown-haired man in a white shirt and loosened striped tie gestures toward the board. Cardboard boxes and overflowing stacks of mail crowd the floor beneath it.
ModelA blue dolphin leaps above a brilliant turquoise sea, its pale belly turned away from the rainbow. A bright rainbow arches behind its back between soft white clouds. The dolphin’s tail curves back toward a little burst of white spray, while gentle blue waves fill the space beneath the rainbow.
SetA masked swordsman dressed entirely in black sits opposite a small balding man at a rough stone table on a rocky hillside. Two silver goblets stand between them. Behind the balding man, a blindfolded blonde woman in a red dress sits with her hands bound. The swordsman rests one gloved hand near his goblet.
SetA gigantic blue toy locomotive with a round gray smiling face bursts through the side of a two-story suburban house. The locomotive is nearly as tall as the house, and its red buffer beam projects over the front lawn. Broken wall panels lie below it. Two people beside a small parked car look up at the engine.
SetA blonde woman in a black off-the-shoulder top leans across a restaurant table, shouting and pointing directly at a white cat seated opposite her. A dark-haired woman stands behind her, gripping her shoulders and trying to pull her back. The cat sits upright on a dining chair with its ears splayed and its mouth slightly open, looking back at the pointing woman. A plate of green salad sits directly in front of the cat. Drinking glasses, cutlery, and folded napkins cover the table, with upholstered chairs and a low restaurant partition surrounding the confrontation.
SetA huge, roughly spherical ball is covered in household objects pointing in every direction: a chair, a television, an umbrella, a frying pan, a banana, and several books. A tiny green figure with a long cylindrical head pushes against its base. Ahead of the ball, a traffic cone and a teapot remain on the floor.
Alt-BuildA green sea serpent rises from the water beside a small pirate ship. Its neck bends over the bow, while its tail emerges from the water behind the stern.
Alt-BuildA tiny blue-armored knight braces behind a shield and points a lance at a snail twice his height. The snail has a large spiral shell and two raised eye stalks. They face each other across a grassy patch.
Alt-BuildA steam locomotive crossing a bridge over a river. The bridge has a single arch directly beneath the engine.

The ceilingA floor a simulator can verify, an open ceiling on design quality

Physical validity

Connected? Collision-free? Stable under gravity?
→ pass or fail

Design quality

Clever part usage, shapes, proportions, and details.
→ no definitive pass/fail test

We asked people familiar with LEGO design to tell agent builds from human-designed models of a similar part count, as a sort of "LEGO Turing test."

Human-designed models from the BrickNet dataset

How often raters took each system’s build for the human-designed one; 0.5 is chance.

Human-designed models from the BrickNet dataset.

Raters consistently found the human build. No agent was taken for human more than one time in five.

What remainsAgents learn to build what works, but not yet what makes a great design

Design quality

Part usage, proportion, and surface detail.
Headroom lies here

Physical validity and prompt adherence

Connections, collisions, stability, and the prompt.
Largely satisfied with BrickAgent

Paper

The paper's pages

Abstract

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch.

BibTeX

@misc{kulits2026brickbenchevaluatingagenticbrick,
      title={BrickBench: Evaluating Agentic Brick Design},
      author={Peter Kulits and Yiqing Xu and R. Kenny Jones and Cordelia Schmid and Jiajun Wu},
      year={2026},
      eprint={2610.12452},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2610.12452},
}