Toggle light / dark theme

3: Our Most Powerful Foundation World Model

Odyssey-3: a new step toward world models for physical AI

Odyssey has unveiled Odyssey-3, its latest foundation world model, designed to learn how objects move, interact, and respond to actions over time. Unlike conventional video generation that primarily produces visual sequences, Odyssey-3 aims to generate interactive environments that evolve in response to human or AI actions, potentially providing a more useful foundation for physical AI.

Built as an autoregressive diffusion transformer, the model learns patterns of physics, dynamics, and cause-and-effect from video, annotated events, gameplay, and simulated physical interactions. Its applications include generating real-time environments, creating training grounds for AI agents, and adapting learned representations to control physical systems.

Odyssey reports that Odyssey-3 Pro achieved a score of 66.1 on the Physics-IQ Verified video-to-video benchmark, which the company describes as a state-of-the-art result. Its evaluations on WorldMark also placed it first in three of four environment categories. In demonstrations, policies built using Odyssey-3 have been applied to robot-arm manipulation, humanoid tasks, and autonomous driving. The company reports that its driving policy was trained using just 20 hours of driving data while keeping the model’s backbone frozen.

The broader ambition is to move world models beyond visual prediction toward systems that can help machines anticipate how their environments change and learn how to act within them. If these capabilities generalize reliably beyond demonstrations and benchmarks, world models could become useful infrastructure for robotics, autonomous systems, and training increasingly capable AI agents.

The important caveat: generating physically plausible video is not the same as possessing a complete or accurate model of the real world. Benchmark results and demonstrations are promising, but robust transfer to unfamiliar conditions, reliable long-horizon predictions, and safe real-world control still require independent validation.

#worldmodels #robotics #AutonomousSystems #ArtificialIntelligence


Today we’re launching Odyssey-3, our most powerful foundation world model yet. Odyssey-3 Pro sets a new state of the art on Physics-IQ Verified’s benchmark, and ranks 1st in 3 of WorldMark’s 4 categories in our evaluations. Our research preview is available now, and if you’re a physical AI developer interested in building with Odyssey-3, please get in touch.

Odyssey-3 is a learned dynamical system, implemented as an autoregressive diffusion transformer, that predicts how objects move and interact through space and how situations evolve over time. It learns representations of physics, dynamics, and cause-and-effect from a broad dataset of visual observations, and developers use that knowledge both to simulate environments and to train policies for different physical systems.

Our founding team spent a decade building driverless cars, where predicting the world was essential to determining what a car should do next. We founded Odyssey to pursue that idea far beyond the roads, to build a general-purpose technology that could bring learned world knowledge to all machines and tasks. With Odyssey-3, we believe this idea is now being realized.

Leave a Comment

Lifeboat Foundation respects your privacy! Your email address will not be published.

/* */