Aditya Ramabadran, Simon Mahns and Tobias Gessler, engineers at the AI startup Axiom, decided to test a large language model in a real-world setting. They chose a 2024 Toyota Corolla parked in the drive-thru lane of an In-N-Out Burger in the Bay Area and asked OpenAI’s GPT-6 Astra to take control of the vehicle for a quick lunch run.
The trio mounted a laptop in the car and linked a chat interface to a backend server that received feeds from several windshield-mounted cameras and could command the power-steering actuator. A human safety driver remained in the seat with a foot lightly on the brake, ready to intervene if the AI’s commands threatened to breach safe operation.
GPT-6 Astra, a model primarily built for generating text, code and occasional images, managed to steer the Corolla forward, align it with the ordering window, and stop at the pick-up point without any pre-programmed driving instructions. The engineers joked that the episode hinted at artificial general intelligence, noting that traditional autonomous vehicles rely on purpose-built perception and planning stacks, whereas this demonstration leveraged a general-purpose language system.
The experiment aligns with a growing interest in measuring how language models comprehend physical environments. Companies such as Elorian and Scale AI have introduced a benchmark called Humanity’s Sixth Sense to evaluate a model’s ability to interpret real-world scenes, while the Axiom engineers created their own DrivingBench test to gauge navigation performance on a simple parking-lot course.
Before settling on GPT-6, the trio experimented with SpaceXAI’s Grok, which initially refused to issue motion commands, responding that it could only interpret road images. After careful prompting, however, Grok and later models from OpenAI and Anthropic began to produce steering directives, suggesting that multimodal training on images, video and 3D data can give rise to unexpected vehicular capabilities without explicit driving data.
DrivingBench results show that only Astra managed to complete the course, and even then at a crawl, while Claude Fable 5.1 covered roughly 45 percent and Grok only 11 percent of the loop. The engineers observed that the models appeared to adjust their commands in real time, learning from mistakes through in-context feedback, a behavior that could foreshadow more capable physical reasoning but also raises safety concerns for future deployments.