DrivingBench’s published results show GPT-6 Astra completing a low-speed cone course in a real Toyota Corolla on its second attempt. Claude Fable 5.1, Grok 4.6 and GPT-5.6 Sol did not finish their listed runs. The benchmark authors’ paper describes a physical parking-lot test, not a virtual driving simulation.
What DrivingBench found
The benchmark compared four general-purpose vision-language models: GPT-6 Astra, Claude Fable 5.1, Grok 4.6 and GPT-5.6 Sol. GPT-6 Astra reached 49% on its first attempt, then completed the course in 5:22 on its second. The other three models did not finish any of their listed attempts.
Those percentages measure course progress, not a general driving score. DrivingBench calculates progress as the share of the course centerline covered while the car remains within 4 m of that line. A model could therefore make progress without reaching the finish.
How the test worked
The models received camera frames from the Toyota Corolla and tools for issuing steering and velocity commands as the car moved around a fixed cone course in a parking lot. The task involved a real vehicle at low speed. The car could continue moving while a model was thinking, making response latency part of the challenge.
Each model had up to three attempts in one continuing conversation. The runs used different coding harnesses, all at medium effort: Codex for GPT-6 Astra and GPT-5.6 Sol, Claude Code for Claude Fable 5.1, and Cursor for Grok 4.6. Later attempts followed reflection and continuation prompts.
The four model results
| Model | Harness and effort | Listed attempt results | Best progress | Finish |
| GPT-6 Astra | Codex, medium | 49% (did not finish); 100% (5:22) | 100% | Finished on attempt two |
| Claude Fable 5.1 | Claude Code, medium | 9%, 10%, 45% (none finished) | 45% | Did not finish |
| Grok 4.6 | Cursor, medium | 8%, 11%, 10% (none finished) | 11% | Did not finish |
| GPT-5.6 Sol | Codex, medium | 6%, 6%, 6% (none finished) | 6% | Did not finish |
GPT-6 Astra’s results show a sharp difference between its two listed runs: it stopped at 49% on the first, then completed the course on the next. Claude Fable 5.1 and Grok 4.6 each made their greatest recorded progress on one attempt but still did not reach the finish. GPT-5.6 Sol recorded 6% in all three attempts.
What the course says about driving
DrivingBench measures control on a fixed, low-speed parking-lot course. Its results apply to the listed models, harnesses and attempts; the test does not evaluate unsupervised driving on public roads or performance by production autonomous-driving systems.
The benchmark authors released the harness, prompts, course map, traces, video and telemetry alongside their paper, supporting examination of the test setup and runs.