A growing consensus across the artificial intelligence sector suggests that large language models will eventually hit a wall imposed by their fundamental inability to interact with the physical environment. Because they are trained exclusively on text, LLMs lack the innate spatial instincts needed to steer robotic limbs, guide autonomous vehicles, or handle physical tasks requiring delicate precision. To resolve this structural limitation, prominent computer scientists including Fei-Fei Li and Yann LeCun have turned their focus toward an alternative architecture: world models.
Developing an intuitive grasp of physical dynamics demands that world models digest vast combinations of visual streams synchronized with physical actions. Before a robotic arm can reliably position components along an assembly line, the governing model must study high-definition floor footage alongside precise telemetry detailing grip firmness, rotational torque, and contact resistance. While LLMs consumed massive libraries of written language scraped across the web, developers of world models face a barren landscape with no equivalent repository of cause-and-effect physical interactions.
The Scarcity of Physical Cause and Consequence
Xiatian Zhu, an associate professor specializing in AI at the University of Surrey, points out that the fundamental pillar of world models is mastering cause and consequence. The public internet contains virtually none of this granular operational data, leaving developers without the foundational material required to train systems on how objects behave when pushed, lifted, dropped, or balanced.
Worldmodeldata, a British startup operating with Yann LeCun as an advisor, seeks to unlock this data bottleneck by aggregating the immense exhaust generated by commercial video game platforms. When gamers play, their systems record continuous controller actions alongside shifting three-dimensional perspectives. While enterprises such as General Intuition and Niantic collect gameplay data directly from their proprietary gaming environments, Worldmodeldata acts as a centralized broker. The company organizes and licenses this material so research facilities can bypass negotiating separate data-sharing pacts with dozens of individual game development houses.
Converting Gaming Experiences into Machine Intelligence
Rhea Loucas, the chief executive officer of Worldmodeldata, argues that millions of advanced video games exist today, and their simulated worlds increasingly mimic the real one. From her perspective, drawing upon the deep, diverse, and abundant scenarios created within video games represents a logical pathway to educating advanced artificial intelligence.
Researchers operate under the prevailing assumption that world models will scale in performance as their training datasets expand, replicating the developmental trajectory seen in text-based language models. Because empirical validation of this scaling law remains an active effort, the sheer scarcity of high-fidelity physical data represents the single largest operational bottleneck for AI labs attempting to build embodied systems.
Previous efforts to generate physical interaction data relied on placing tracking sensors on human operators and robotic hardware inside controlled testing rooms. This laborious process yields only small trickles of usable telemetry and completely misses the chaotic, unexpected scenarios encountered in unstructured outside environments. Nicole Fraenkel, a partner at Khosla Ventures and an investor in General Intuition, explains that while workers can be hired to demonstrate repetitive pick-and-place tasks, repetition by itself fails to capture the true disorder of the real-world spaces where machines must eventually operate.
Conquering Corner Cases and Monetizing Play
The operational premise behind Worldmodeldata rests on the conviction that video game logs offer the exact scale and variety required to catalog critical fringe scenarios. Inside gaming engines, 3D coordinate mapping is continuously paired with instantaneous user inputs, capturing unexpected trajectories and rare interactions that physical labs cannot safely simulate.
Fraenkel highlights that mastering these corner cases is the central challenge in deploying autonomous hardware. When deploying an autonomous car, commercial drone, passenger plane, warehouse forklift, or four-legged robotic platform, the financial and physical cost of a perceptual error is exceptionally severe.
Worldmodeldata has already secured licenses covering nearly 1 million hours of gameplay records from studios producing well-known commercial titles, though Loucas chose not to disclose the identities of those developers. Looking ahead, the startup intends to build payment pipelines so individual players can receive compensation for sharing their gaming telemetry. Loucas anticipates that gameplay records will ultimately constitute the vast majority of core training datasets for world models, which engineers can then calibrate with bespoke sensory feeds tailored to specific real-world tasks. In her assessment, this strategy could deliver a breakthrough milestone for world models comparable to the arrival of generative language tools.
Nvidia and Skeptics Urge Caution Over Video Game Physics
The enthusiasm surrounding video game telemetry is not universal across the tech landscape. Nvidia, which produces dedicated world models engineered to function on its specialized semiconductors, bypasses commercial gaming inputs in favor of proprietary simulation engines designed specifically to calculate real-world physical laws with mathematical rigor.
Ming-Yu Liu, who directs world model engineering at Nvidia, cautions that models trained on gaming telemetry will struggle when deployed on physical operations requiring delicate motor dexterity. Video game software depends heavily on visual approximations and programmatic shortcuts to create the illusion of real movement without computing the actual underlying forces. For example, a character on screen may appear to pick up an apple from a table, but the engine does not model the exact distribution of frictional pressure applied by individual fingertips to keep the object stable.
Liu advises taking a conservative stance on applying gameplay telemetry to physical manipulation tasks, noting that the underlying physics of material handling is vastly more complicated. He suggests that gaming data is better suited for world models that generate photorealistic video sequences or synthetic 3D landscapes. Zhu echoes this skepticism, observing that while video games function as rudimentary simulators with a degree of physical reference, their representations remain coarse and approximate. Nonetheless, as researchers pursue the ultimate breakthrough in embodied AI, diverse methodologies remain under active evaluation, and Fraenkel notes that definitive judgment on which training architecture will succeed remains unsettled.



















