In DrivingBench’s AI-agent driving test, GPT-6 Astra was the only one of four listed configurations to finish the short, cone-marked parking-lot course. It reached 100% on its second attempt in 5:22. Across 11 listed attempts, eight reached 11% progress or less.

DrivingBench’s short parking-lot course

The benchmark put general-purpose AI configurations on a short route marked with cones and curves. DrivingBench lists up to three attempts per configuration in one continuous chat. Its entries identify Codex as the interface for GPT-6 Astra and GPT-5.6 Sol, Claude Code for Claude Fable 5.1, and Cursor for Grok 4.6; each entry carries the label “medium.”

Four configurations, 11 attempts

GPT-6 Astra was the only listed configuration to finish DrivingBench’s course

DrivingBench’s leaderboard records the following progress and finish results:

Model configurationAttempt progressBest progressFinished the course
GPT-6 Astra49%, 100%100%Yes, on attempt 2
Claude Fable 5.19%, 10%, 45%45%No
Grok 4.68%, 11%, 10%11%No
GPT-5.6 Sol6%, 6%, 6%6%No

Grok 4.6’s three listed attempts all ended without a finish. Its first reached 8% after two accepted commands.

How DrivingBench measured progress

DrivingBench calculates progress from the distance traveled along the course centerline while the vehicle remains within 4 meters of it. The distance counts as a percentage of the centerline route to the finish zone, and progress reached before a collision remains in the total.

The results describe performance on this short parking-lot route; they do not establish how any configuration would perform on public roads.