OpenAI GPT-6 Astra Steers Toyota Corolla Through Drive-Thru Lane in Novel Driving TestAI
8 Oct 2026, 8:38 pm (59 min ago)· 1

OpenAI GPT-6 Astra Steers Toyota Corolla Through Drive-Thru Lane in Novel Driving Test

Engineers connected OpenAI's GPT-6 Astra directly to a car's steering mechanism using camera feeds, successfully directing the vehicle through a fast-food drive-thru lane.

Parked in the drive-thru lane of an In-N-Out restaurant in the Bay Area, three engineers decided against driving their 2024 Toyota Corolla themselves. Opening a laptop, they handed control of the steering wheel to OpenAI's GPT-6 Astra. The setup linked a chat interface to a central server, which processed visual input from several windshield-mounted cameras and relayed steering commands directly into the car's power steering system. To ensure safety throughout the experiment, a driver kept one foot hovering directly above the brake pedal. Designed primarily to craft sentences, software code, and visual images, the language model steadily piloted the vehicle forward until it reached the take-out window where the crew picked up their meal.

Following the successful drive-thru maneuver, one engineer observed that artificial general intelligence, widely known as AGI, might finally be materializing. Autonomous vehicles usually rely on custom-tailored algorithms engineered exclusively for navigation and driving tasks. In contrast, this vehicle was directed entirely on the fly by an AI architecture conceived for text processing, operating without any targeted prior coaching from the engineers. The drive-thru journey demonstrates that language-centric AI models are beginning to grasp an elementary comprehension of the physical world around them.

Also read

Language Models Venture Into the Physical World

Entrusting two tons of moving machinery to a general-purpose language model carries substantial risks, even though the brief lunch trip concluded without incident. Modern AI systems solve complex analytical problems and execute digital workflows effortlessly, yet their capabilities typically collapse outside computer environments and internet servers. Once introduced to physical surroundings, advanced models frequently become disoriented. Despite frequent declarations celebrating the dawn of AGI, leading artificial intelligence organizations continue to treat real-world physical reasoning as an unresolved frontier.

Recognizing this technical void, several researchers have departed established tech corporations to launch specialized ventures targeting physical reasoning. Andrew Dai, formerly a researcher at Google DeepMind and currently the chief executive officer of Elorian AI, notes that enhanced visual reasoning will unlock practical robotic capabilities. These potential applications include software that can assess whether restaurant guests enjoy their meals or autonomous machines managing household duties. Dai emphasizes that physical reasoning represents an indispensable cornerstone for modern home robotics.

Assessing Models on Humanity's Sixth Sense and DrivingBench

To quantify these capabilities, Elorian partnered with Scale AI to introduce a novel evaluation standard called Humanity's Sixth Sense, which grades a model's aptitude for interpreting real physical environments. Xingang Guo, a research scientist at Scale AI who helped design the framework, explained that conventional visual research traditionally prioritized simple perception. The new metric focuses instead on whether an algorithm can intuitively grasp an environment the way a human mind does naturally.

The In-N-Out experiment was conceived and executed outside working hours by Ramabadran, Mahns, and Gessler, who are employed by Axiom. Their primary objective was measuring how effectively language models cope with unscripted real-world friction. The concept originated during a casual weekend gathering after the trio observed Astra producing intricate 3D environment simulations. Living across the Bay Area surrounded by autonomous fleets from Tesla and Waymo, they questioned whether digital 3D spatial competence could translate into real-world vehicle navigation.

Prompting Strategies and Navigational Benchmarks

Initial trials involved SpaceXAI's Grok before expanding to the latest architectures from OpenAI and Anthropic. At first, the models explicitly rejected motion requests, stating that they could analyze road photography but lacked authority to dispatch movement commands to a physical automobile. Through iterative and careful prompt engineering, the team coaxed the models into accepting vehicle oversight, leaving the engineers stunned when the models began executing genuine steering maneuvers.

While major technology labs are aggressively expanding spatial reasoning parameters, the engineers doubt these firms are intentionally training foundational models to drive cars. Instead, driving aptitudes appear to emerge organically as an unintended byproduct of multimodal scaling that ingests video, images, and 3D datasets. Mahns characterized the phenomenon as an emergent behavior stemming directly from larger multimodal training pipelines.

Despite this progress, the trio's standardized vehicular evaluation, DrivingBench, proves that current models remain far from earning a standard driver's license. The benchmark tracks performance around a designated parking-lot driving circuit. Astra stood out as the sole model capable of finishing the course, moving at an exceptionally sluggish pace. Claude Fable 5.1 completed 45 percent of the track, while Grok stalled after covering just 11 percent of the layout. Ramabadran highlighted that the latest models demonstrated an ability to correct course dynamically, showing in-context learning by adjusting to vehicle controls based on operational missteps.

Questions & Answers

Which AI model was used by the engineers to steer the car?
The engineers used OpenAI's GPT-6 Astra, connecting its chat interface through a server to cameras and the car's steering mechanism.
What vehicle was used for the drive-thru test?
The team conducted the experiment using a 2024 Toyota Corolla equipped with windshield-mounted cameras.
How did competing AI models perform on the DrivingBench test?
Astra was the only model to finish the parking lot course, while Claude Fable 5.1 completed 45 percent and Grok achieved only 11 percent.
Were the models intentionally trained to operate automobiles?
No, the models received no driving-specific coaching; the ability emerged organically from extensive multimodal and 3D spatial training.

Comments 0

No comments yet — be the first.

Citizen journalism

Become a TrendKia journalist

Voice of the people

Share news, photos and videos from your area with TrendKia and let your voice reach the nation. Every citizen a journalist.

Join now
CH 01 LIVE
TrendKia TV ON AIR
Chamar no WhatsApp