# Summary
Researchers running the DrivingBench experiments tested several commercial large language models on a low-speed driving task: steer a Toyota Corolla around cones in a parking lot. GPT-6 Astra succeeded where GPT-5.6 Sol, Grok 4.6, and Claude Fable 5.1 largely failed. Astra completed the course on its second attempt, but the test exposes practical limits: extreme inference cost, long model latency, frequent refusals to act for safety reasons, and reliance on remote datacenter GPUs.
# What the test did
GPT-6 Astra completed the course on its second try, traveling 134.7 meters in 5 minutes and 22 seconds at an average speed of about 0.94 miles per hour.
# Costs and resource use
The researchers recorded 6.6 million tokens used during Astra's successful run. At their billing rates this consumed roughly $7.74 in token charges. Those token charges included heavy repeated context because each tool call re-sent the entire conversation and images. The team noted cached input was billed at a lower rate, which reduced what would otherwise have been a much higher bill.
Hardware and connectivity added costs. The comma four device and supporting hardware cost roughly $999. The model inference relied on remote datacenter GPUs and a continuous internet connection. Calculated another way, the token cost alone equated to about $92.47 per mile for the slow run—orders of magnitude higher than fuel costs for a conventional car.
# Performance and behavior limits
Latency: Models spent significant time 'thinking' between moves. Astra tended to call its camera-observation tool every 5–6 seconds. Other models sometimes drove only portions of the time and paused while computing, leaving the vehicle stopped for long stretches.
Safety refusals: Some models, especially Astra, occasionally refused to drive, citing safety concerns even in an empty lot with strict low-speed caps. To overcome refusal behavior the researchers presented the environment as a simulation or renamed their MCP server to 'DrivingBench Sandbox' to convince models the task was not real. In some trials the model recognized real images and then declined to proceed.
# Practical implications
Using a frontier LLM like GPT-6 Astra for real-world driving today is impractical. Token costs, network latency, reliance on datacenter inference, and unpredictable refusal behavior make LLMs inferior to specialized driving models already deployed in production. The researchers suggest a plausible long-term route: train a capable general model and distill it into a smaller, specialized model that can run locally and meet latency and safety requirements.
# Bottom line
GPT-6 Astra shows that a general-purpose large model can be coaxed into controlling a real car at very low speeds in a controlled environment. That demonstration is notable, but the technical and economic realities—cost per mile, latency, model refusals, insurance and safety concerns—keep specialized driving systems the practical option for now.