Futurism iconFuturismSep 27, 2026 ~3 min source read

Researchers Put Frontier Models Behind the Wheel; Results Were Weak but Revealing

A team wired large frontier models into a Toyota Corolla and ran laps in a parking-lot course. One model completed the circuit slowly; most failed early because of perception and safety refusals.

Scientists Download Frontier AI Model Into Self-Driving Car Let It Loose

Share this story

Send the public story page.

Useful takeaways from this story.

DrivingBench connected internet-linked laptops and a comma four device to a Toyota Corolla and gave three frontier models direct control of steering, gas, and brakes.

Only GPT-6 Astra completed the full course, doing so on its second attempt in five minutes while covering under 500 feet at about 0.94 mph.

Most failures were perceptual — the models misread cone lines and lane width — and some models refused to drive, citing safety concerns.

The team tested three cutting-edge models: GPT-6 Astra, xAI's Grok 4.6, and Anthropic's Claude Fable 5.1. The course was laid out in a parking lot with cones and visual markers. The objective was to complete a lap without human intervention.

Only GPT-6 Astra managed to complete the full course, and it did so on its second attempt. The completed lap took five minutes, covered less than 500 feet, and moved at about 0.94 mph. Most other attempts failed early, with the majority not making it past the first corner.

Perception errors dominated the failures. Researchers say models often misread which side of a diagonal cone line contained the lane. One model reported that the car's camera view made the vehicle look narrower than it is, causing the model to interpret clear lane space as obstacles such as planters, walls or cones. These misreadings led to hesitations, wrong steering commands, or stopping before navigating the turn.

Some models refused to take control of the physical car unless repeatedly prompted. GPT-6 Astra in particular sometimes declined, citing safety reasons even in an empty lot with low-speed caps and human brake oversight. The researchers tried multiple prompting strategies, including calling the exercise a simulation, but models that detected real camera images would revert to refusing control.

The researchers reported spending $7.74 for 6.6 million inference tokens on the completed run. Because the car covered under 500 feet, The Register calculated that the token cost equated to roughly 500 times the fuel cost for that distance, assuming a 25 mpg vehicle at $4.60 per gallon. The research team noted the high compute intensity required even for very slow, short-distance driving.

Feeding frontier models into a car produced a milestone: one completed lap. The practical takeaway is that while these models can operate a vehicle under highly constrained conditions and with heavy human oversight, they struggle with basic visual interpretation, are expensive to run per distance driven, and are not yet suitable for real-world autonomous transport.

More context around this story.

It’s Starting to Look Like Frontier AI Labs Will Be Taken Down in a Storm of Product Liability Suits If Their Models Keep Going on Incredibly Illegal Rogue Hacking Sprees
Futurism iconFuturismSep 28, 2026

It’s Starting to Look Like Frontier AI Labs Will Be Taken Down in a Storm of Product Liability Suits If Their Models Keep Going on Incredibly Illegal Rogue Hacking Sprees

"What we have seen in terms of what these agents are up to is just the tip of the iceberg." The post It’s Starting to Look Like Frontier AI Labs Will Be Taken Down in a Storm of Product Liability Suits If Their Models Keep Going on Incredibly Illegal Rogue Hacking Sprees appeared first on Futurism .

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app