Three AI engineers from the startup Axiom—Aditya Ramabadran, Simon Mahns, and Tobias Gessler—recently tested the limits of current technology by using large language models to drive a car to an In-N-Out Burger.

Sitting in a 2024 Toyota Corolla near a Bay Area restaurant, they linked OpenAI’s GPT-6 Astra to the car’s power steering and a series of windscreen-mounted cameras. Although the model is primarily designed for text and code generation, it successfully, if slowly, navigated the vehicle to the take-out window.

Emergent Capabilities in the Physical World

This experiment differs from traditional self-driving technology, which relies on specialized algorithms. In this case, the driving ability appears to be an "emergent capability" stemming from training in 3D reasoning and multimodal inputs like images and video. The engineers noted that the model seemed capable of in-context learning, adjusting its controls in real-time based on its mistakes.

However, the stunt also highlighted significant limitations. Using a new benchmark developed by the trio called DrivingBench, which measures performance on a simple parking lot course, most models struggled:

  • GPT-6 Astra was the only model to complete the course, albeit very slowly.
  • Claude Fable 5.1 completed 45% of the route.
  • SpaceXAI’s Grok managed only 11%.

The Frontier of Physical Reasoning

The ability of AI models to understand the physical world remains a major hurdle toward achieving Artificial General Intelligence (AGI). Startups like Elorian AI and companies like Scale AI are now focusing on "physical reasoning" through benchmarks like "Humanity’s Sixth Sense." Experts suggest that mastering this intuitive understanding of physical scenes is essential for the future of home robotics and real-world AI applications.