In its report published September 23, 2026, Epoch AI said GPT-6 Astra scored 80% on the Furniture Assembly Benchmark (FAB) in September 2026. The test asks models to identify assembly mistakes in photos of partly assembled IKEA furniture; it evaluates visual diagnosis, not physical assembly.
GPT-6 Astra’s reported score on FAB
Epoch AI’s September 2026 results put GPT-6 Astra at 80%. The model’s score varied across the three tested builds: 71% for the STÄLL shoe cabinet, 96% for the TONSTAD bed frame and 67% for the GULLABERG dresser. Its reported median time was three minutes per photo.
What the Furniture Assembly Benchmark tests
FAB uses 60 images covering three IKEA furniture builds, including correctly assembled examples and builds with intentional mistakes. Models receive the assembly instructions, an image-zoom tool and a Python interpreter in a sandbox, with a limit of 80 steps per image.
To score a flawed build, a model must identify every mistaken assembly step and describe the error. For a correctly assembled item, it must recognize that there is no mistake. GPT-5.6 Sol evaluates the descriptions; Epoch AI says it was prompted to grade leniently.
That makes FAB a test of matching visual details against assembly instructions—not a demonstration that GPT-6 Astra, or a robot, can physically put furniture together.
How the reported scores compare over time
Epoch AI reported Claude Opus 4.5 as its previous highest-scoring model on FAB, with 28% in November 2025. The two results come from different benchmark periods.
| Model | Reported FAB score | Benchmark period |
| GPT-6 Astra | 80% | September 2026 |
| Claude Opus 4.5 | 28% | November 2025 |
What the results cover
The evaluation covers 60 images from three furniture builds. Epoch AI says it has not measured human performance on FAB and that it remains unclear how well the results extend to other physical tasks.