A newly reported DrivingBench benchmark found that GPT-6 Astra was the only one of four general-purpose vision-language models to complete a low-speed cone course in a Toyota Corolla. Astra finished on its second attempt. The test was a prepared parking-lot run, not a measure of how these models handle ordinary traffic.
GPT-6 Astra was the only reported finisher
DrivingBench tested GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Grok 4.6. Each model had up to three attempts in one conversation. Astra completed the course on attempt two; no other attempt passed half the course.
The reported figures give more detail about individual runs. Astra covered 134.7 meters in five minutes and 22 seconds on its finishing attempt. Claude Fable 5.1 reached approximately half the course on attempt three. GPT-5.6 Sol covered 17.1 meters—about 6% of the course—on its third attempt, while Grok 4.6 reached 22.6 meters, about 11%, in its best attempt.
How the models controlled the Corolla
The models received camera frames and used three tools to issue steering and velocity commands. In other words, they had to interpret what the camera saw and send instructions directly to the car’s controls.
There was a timing challenge, too: the Corolla could keep moving while a model processed an image, and a new command replaced the one already running. That made the time needed to interpret the scene and respond part of the task, not just the quality of each instruction.
Two of the four models improved materially across attempts when they could retain context. The repeated runs therefore tested performance with accumulated conversational context, as well as a single attempt at steering.
A prepared course is not a public-road test
The benchmark measured performance on a low-speed course laid out with cones in a parking lot. It does not establish how the models would perform amid ordinary traffic or whether they are ready to drive on public roads. A person was reportedly ready to apply the brake during the test.
Researchers released the control harness, prompts, course map, traces, video, and telemetry.