Aditya Ramabadran, Simon Mahns, and Tobias Gessler, three AI engineers at a startup called Axiom, recently used artificial intelligence to drive a 2024 Toyota Corolla to an In-N-Out Burger in the Bay Area. They connected OpenAI's GPT-6 Astra to the car's power steering system and mounted cameras to navigate the vehicle through the drive-thru. A safety driver was present to intervene if necessary.
The AI model, typically used for generating text and code, successfully navigated to the take-out window. One engineer remarked, "Maybe AGI is here after all," referring to the concept of artificial general intelligence.
While self-driving cars are not new, this experiment utilized a model not specifically designed for driving tasks. The engineers noted that the experiment indicates language-based AI models may be developing a basic understanding of the physical world.
Although the drive was uneventful, the engineers acknowledged the risks of using a general-purpose AI model in control of a vehicle. Current AI models excel in virtual tasks but struggle in real-world applications.
Andrew Dai, CEO of Elorian AI and former researcher at Google DeepMind, stated that advancements in visual reasoning could lead to new applications for AI, such as understanding customer experiences in restaurants or functioning robots in homes.
Elorian and Scale AI have developed a benchmark called Humanity’s Sixth Sense to measure AI models' understanding of physical scenes. Xingang Guo, a research scientist at Scale AI, emphasized the goal of assessing a model's intuitive understanding of scenes.
The engineers conceived the idea for their project while socializing and observing the capabilities of AI models like Astra. They initially tested SpaceXAI's Grok but found that the models were hesitant to take control of the vehicle. With prompting, they were able to drive.
The engineers believe that while AI companies are working on spatial reasoning, this does not necessarily mean they are training models for driving. They suggest that the driving capabilities observed may be an emergent property of enhanced multimodal training.
Their new benchmark, DrivingBench, indicates that AI models still have significant limitations in driving skills. Only Astra completed a simple parking lot course, and it did so slowly. Claude Fable 5.1 and Grok completed 45% and 11% of the course, respectively.
Ramabadran noted that the latest models showed some ability to adapt to errors and improve their driving in real time. "The models really did seem to be adjusting or in-context learning based on their mistakes and learning how to better navigate the controls," he said.