With the twirl of a thumbstick and the press of a few buttons, even an unskilled player can navigate a complex 3D environment. A British startup, Worldmodeldata, is wagering that these straightforward sequences contain a trove of information that can be used to train a new class of artificial intelligence.
Beyond Text: The Rise of World Models
There is a growing belief within the AI industry that large language models (LLMs) are fundamentally limited by their inability to navigate the physical world. Trained solely on text, they lack the finesse required to pilot autonomous vehicles or operate robotic arms. To bridge this gap, prominent researchers like Fei-Fei Li and Yann LeCun are focusing on "world models."
To become fluent in real-world physics, these models require a combination of visual and action data—understanding cause and consequence. Unlike the vast oceans of text available for LLMs, there is no comparable corpus for physical interaction. Worldmodeldata aims to solve this by packaging controller inputs and data from video game studios into training datasets.
The Debate Over Digital Physics
The startup’s hypothesis is that video games provide the necessary scale and diversity to capture "corner cases"—unusual scenarios that are critical for the safety of autonomous cars or factory robots. However, industry giants like Nvidia remain skeptical. Ming-Yu Liu, who leads world model development at Nvidia, argues that video game physics are often "eccentric" and rely on shortcuts to create the illusion of realism.
Critics point out that games often lack fine-grained motor control data, such as the specific torque or pressure needed to grip an object. Despite these concerns, Worldmodeldata has already licensed nearly 1 million hours of data from popular titles. The company’s CEO, Rhea Loucas, believes this approach could lead to a "GPT moment" for world models, finally making them useful for real-world physical tasks.