AI NewsModels & agentsAnnouncement
DrivingBench let four frontier models steer a real Toyota Corolla around a cone course, and only GPT-6 Astra finished it
DrivingBench handed the steering wheel, accelerator and brakes of a real Toyota Corolla to four frontier models on a fixed cone course, and only GPT-6 Astra finished, in five minutes twenty-two seconds, on 246 million tokens and $7.74.

Image: DrivingBench
Why it mattersOne successful driving attempt at $7.74 in tokens is a plain number for what it costs today to ask a cloud model to run a physical task step by step, and the 55 to 94 point gap between the top model and the other three is the difference between a run that finishes and a run that stops halfway.
Four frontier models were handed the steering wheel, the accelerator and the brakes of a real Toyota Corolla on a fixed cone course this week, and only one of them finished. GPT-6 Astra reached the finish zone on its second attempt, running for five minutes and twenty-two seconds on 246.6 million tokens at a cost of $7.74. Claude Fable 5.1 got 45 percent of the way through the course at its best, Grok 4.6 stopped at 11 percent, and GPT-5.6 Sol got 6 percent.
The benchmark is at drivingbench.com and was built by Aditya Ramabadran, Simon Mahns and Tobias Gessler. It reached 290 points and 227 comments on Hacker News in eighteen hours.
What is being measured
Each model gets up to three attempts in one continuous chat session. Progress is scored as how far along the course centreline the car reaches while staying within four metres of it, as a share of the length to the finish zone. A run that leaves the corridor stops counting.
The cars operate, in the creators' words, at "extremely low speeds" and with timestamps written into each frame so a model can watch itself and correct. Ramabadran acknowledged in the Hacker News thread that the delay of sending each frame to a cloud model and waiting for a reply is too long for real driving. He described the whole thing as "just sort of a fun benchmark."
What the numbers actually say
The cost is the interesting number. $7.74 and 246 million tokens is the price of one successful attempt for one model on a course a person can drive in a few minutes without thinking. It is a plain number for what it costs today to ask a cloud model to run a physical task step by step in the real world, before any repeated attempts and before running more than one car.
The gap between the models is bigger than the four scores suggest on first reading. GPT-6 Astra reached 100 percent. The next-best model, Claude Fable 5.1, reached 45 percent, less than half the course. The other two never reached double figures. If a task requires a full completion rather than a best-effort attempt, the difference between one model and the others here is the difference between a run that works and a run that stops halfway.
The sample is one course, one car, and up to three attempts per model, and the benchmark's own authors ran it, so the ranking is only as trustworthy as that. It is enough to say that frontier models can now steer well enough to complete this specific task some of the time, and one is much better at it than the others.
Source
- Benchmark: DrivingBench
- Discussion, including author responses on latency: Hacker News
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


