In DrivingBench’s AI-agent driving test, GPT-6 Astra was the only one of four listed configurations to finish the short, cone-marked parking-lot course. It reached 100% on its second attempt in 5:22. Across 11 listed attempts, eight reached 11% progress or less.
DrivingBench’s short parking-lot course
The benchmark put general-purpose AI configurations on a short route marked with cones and curves. DrivingBench lists up to three attempts per configuration in one continuous chat. Its entries identify Codex as the interface for GPT-6 Astra and GPT-5.6 Sol, Claude Code for Claude Fable 5.1, and Cursor for Grok 4.6; each entry carries the label “medium.”
Four configurations, 11 attempts
DrivingBench’s leaderboard records the following progress and finish results:
| Model configuration | Attempt progress | Best progress | Finished the course |
| GPT-6 Astra | 49%, 100% | 100% | Yes, on attempt 2 |
| Claude Fable 5.1 | 9%, 10%, 45% | 45% | No |
| Grok 4.6 | 8%, 11%, 10% | 11% | No |
| GPT-5.6 Sol | 6%, 6%, 6% | 6% | No |
Grok 4.6’s three listed attempts all ended without a finish. Its first reached 8% after two accepted commands.
How DrivingBench measured progress
DrivingBench calculates progress from the distance traveled along the course centerline while the vehicle remains within 4 meters of it. The distance counts as a percentage of the centerline route to the finish zone, and progress reached before a collision remains in the total.
The results describe performance on this short parking-lot route; they do not establish how any configuration would perform on public roads.