GPT-6 Astra Finished a Parking-Lot Drive Test. Three AI Models Did Not.

Only GPT-6 Astra finished a 135-meter cone course at under 1 mph; three rival models failed before the first turn

Al Landes Avatar
Al Landes Avatar

By

Image: Drivingbench

Key Takeaways

Key Takeaways

  • GPT-6 Astra alone finished a 135-meter cone course; three rival AI models failed early.
  • Latency transformed processing delays into physical hazards, exposing AI driving’s core real-world vulnerability.
  • Spatial reasoning failures caused models to misread cone gaps and invent nonexistent color-coding rules.

Researchers wired a 2022 Toyota Corolla to Comma Four hardware — a third-party driver-assistance harness built on OpenPilot — and handed the controls to the same AI models answering your emails. No simulation. A real car, moving, in an empty parking lot. Four models entered. Most runs failed before clearing the first turn of a 135-meter cone course. One model spotted a gap between boundary cones and treated it as an open door. This is not a story about robotaxis arriving. It’s a story about how unforgiving the physical world is.

What “Driving” Actually Meant Here

Three tools, one moving car, and a latency problem measured in feet.

Researchers Aditya Ramabadran, Simon Mahns, and Tobias Gessler gave each model three functions: observe, move, stop. One command at a time. Camera frames and vehicle telemetry fed back between each action. The catch: the car kept moving while the model processed its next response. Latency wasn’t a UX inconvenience — it was a physical hazard. A human supervisor sat ready to brake throughout. This wasn’t Tesla Autopilot. It was asking a chatbot to think like a driver, one deliberate command at a time.

By the numbers — how each model performed:

  • GPT-6 Astra: Only finisher. Completed approximately 134.7 meters on its second attempt in 5 minutes 22 seconds at roughly 0.94 mph. The successful run consumed approximately 6.6 million tokens at a reported cost of $7.74 in inference charges, according to the benchmark’s published report.
  • Claude Fable 5.1: Reached approximately 45% of the course. Did not finish.
  • Grok 4.6: Reached approximately 11% at best. First attempt ended after two commands.
  • GPT-5.6 Sol: Approximately 6% across all attempts. Failed to clear the opening section.

When the Model Invents Its Own Rules

Spatial reasoning broke first — and it broke fast.

Grok’s first run ended because it read a gap between boundary cones as a gate and drove toward it. It’s the AI equivalent of someone confidently pushing a clearly marked “STAFF ONLY” door at a concert venue — the information was there. The interpretation failed completely. Separately, another model reportedly invented a color-coding rule for cones despite explicit instructions stating the colors weren’t uniform. These failures point to the same core weaknesses: spatial reasoning, instruction adherence, and feedback control all proved unreliable under real-world conditions.

Astra’s win deserves equal honesty. According to the benchmark’s published results, it hugged one side of a long bend and nearly left the course before reaching the finish box. A victory the way finishing a 5K with one shoe counts — it still counts.

The broader result from the benchmark’s 11 recorded runs is consistent: only Astra’s second attempt and Claude’s third progressed meaningfully beyond the first corner, and no attempt other than Astra’s second passed 50% of the course.

What One Completed Lap Actually Proves

A parking-lot loop is not a school run on a wet Tuesday.

Astra crossing that finish box is real and genuinely notable. It is also 0.94 mph on a supervised 135-meter course with a human ready to brake at any moment. DrivingBench describes the project as an open benchmark with recorded traces — not a demonstration that these models are ready for public roads. The gap between a slow, controlled parking-lot loop and consistent, reliable driving in real traffic is not a product cycle away. One completed lap does not close it.

Share this

Every news piece, car review, and list is fueled by real human research and experience. See how we keep it real in our Code of Ethics →


Al Landes Avatar