Researchers wired a 2022 Toyota Corolla to Comma Four hardware — a third-party driver-assistance harness built on OpenPilot — and handed the controls to the same AI models answering your emails. No simulation. A real car, moving, in an empty parking lot. Four models entered. Most runs failed before clearing the first turn of a 135-meter cone course. One model spotted a gap between boundary cones and treated it as an open door. This is not a story about robotaxis arriving. It’s a story about how unforgiving the physical world is.
What “Driving” Actually Meant Here
Three tools, one moving car, and a latency problem measured in feet.
Researchers Aditya Ramabadran, Simon Mahns, and Tobias Gessler gave each model three functions: observe, move, stop. One command at a time. Camera frames and vehicle telemetry fed back between each action. The catch: the car kept moving while the model processed its next response. Latency wasn’t a UX inconvenience — it was a physical hazard. A human supervisor sat ready to brake throughout. This wasn’t Tesla Autopilot. It was asking a chatbot to think like a driver, one deliberate command at a time.
By the numbers — how each model performed:
- GPT-6 Astra: Only finisher. Completed approximately 134.7 meters on its second attempt in 5 minutes 22 seconds at roughly 0.94 mph. The successful run consumed approximately 6.6 million tokens at a reported cost of $7.74 in inference charges, according to the benchmark’s published report.
- Claude Fable 5.1: Reached approximately 45% of the course. Did not finish.
- Grok 4.6: Reached approximately 11% at best. First attempt ended after two commands.
- GPT-5.6 Sol: Approximately 6% across all attempts. Failed to clear the opening section.
When the Model Invents Its Own Rules
Spatial reasoning broke first — and it broke fast.
Grok’s first run ended because it read a gap between boundary cones as a gate and drove toward it. It’s the AI equivalent of someone confidently pushing a clearly marked “STAFF ONLY” door at a concert venue — the information was there. The interpretation failed completely. Separately, another model reportedly invented a color-coding rule for cones despite explicit instructions stating the colors weren’t uniform. These failures point to the same core weaknesses: spatial reasoning, instruction adherence, and feedback control all proved unreliable under real-world conditions.
Astra’s win deserves equal honesty. According to the benchmark’s published results, it hugged one side of a long bend and nearly left the course before reaching the finish box. A victory the way finishing a 5K with one shoe counts — it still counts.
The broader result from the benchmark’s 11 recorded runs is consistent: only Astra’s second attempt and Claude’s third progressed meaningfully beyond the first corner, and no attempt other than Astra’s second passed 50% of the course.
What One Completed Lap Actually Proves
A parking-lot loop is not a school run on a wet Tuesday.
Astra crossing that finish box is real and genuinely notable. It is also 0.94 mph on a supervised 135-meter course with a human ready to brake at any moment. DrivingBench describes the project as an open benchmark with recorded traces — not a demonstration that these models are ready for public roads. The gap between a slow, controlled parking-lot loop and consistent, reliable driving in real traffic is not a product cycle away. One completed lap does not close it.
























