model release

Odyssey-3 world model enters free public preview; 14B base model, 66.1 Physics-IQ claim carries caveats

TL;DR

Odyssey has opened a free public research preview of Odyssey-3, an autoregressive diffusion transformer that generates interactive environments from text prompts in real time. The base model has 14 billion parameters and outputs 832 × 480 video. Odyssey's headline Physics-IQ score of 66.1 does not meet the benchmark's own requirements for a record claim.

4 min read
0

California-based Odyssey has opened a free public research preview of Odyssey-3, a world model that generates interactive environments from text prompts in real time. The base model has 14 billion parameters and generates video at 832 × 480 pixels. Odyssey-3 Pro supports 1280 × 720. Developers can apply for API access. Pricing is not yet disclosed.

Founders Oliver Cameron and Jeff Hawke first unveiled the model on September 15, with a focus on robotics, autonomous driving and video games. The new release adds public access, technical detail and benchmark results.

Architecture and the demo

Odyssey-3 is an autoregressive diffusion transformer. It continuously generates new video frames conditioned on previous frames and user actions. According to Odyssey, the model learns physical relationships and cause and effect from visual observations. Training data included:

  • Internet videos with event descriptions
  • Video game footage paired with keyboard and mouse inputs
  • Simulated physical interactions

Odyssey says an additional training technique reduces the number of compute steps required, which makes real-time generation possible. The free online demo runs on Odyssey-3 Flash. Users can pick first-person or third-person perspectives, move through the generated world and trigger events. Context window, training cutoff and Flash's parameter count are not disclosed.

Benchmarks: the caveats matter

Odyssey says Odyssey-3 Pro scores 66.1 on the video-to-video Physics-IQ Verified benchmark. The test covers fluid mechanics, optics, solid mechanics, magnetism and thermodynamics. Models continue videos of real-world experiments, and the outputs are compared with actual outcomes.

The 66.1 figure comes from a single run in which a selection method picked one of eight generated videos per task. Physics-IQ rules require four runs with standard deviation reported for any record claim, so the score does not meet that bar. Without selection, Odyssey-3 Pro averaged 63.37 across four runs. Both results are on the official leaderboard, but Odyssey submitted them itself.

On WorldMark, which tests control-following, visual quality and long-term consistency, Odyssey's own evaluation shows:

Category Score Rank
First-Person Stylized 77.2 1st
Third-Person Real 79.0 1st
Third-Person Stylized 76.3 1st
First-Person Real 80.6 3rd

These results have not been independently verified.

Beyond generation: controllers

Odyssey positions the model as a foundation for controlling other systems. Each application pairs the world model with a specialized controller that translates predictions into commands. All of the following are company claims:

  • Robotic arms: A few dozen hours of demonstration data was enough to train a controller for multiple arms. The robots recovered from failed grasps that were not in the training data.
  • Humanoids: Odyssey is working with Swiss robotics company Flexion. Controllers Flexion built on Odyssey-3 reportedly perform more reliably than comparison models when conditions change.
  • Drones: A controller trained on simulated flight data dodged obstacles and flew to targets in a virtual indoor environment.
  • Games: A controller trained on roughly two hours of GTA V footage drove vehicles and fought enemies. It transferred to Red Dead Redemption 2 without extra training.

Odyssey also demoed an agent that received a natural-language task and acted inside an Odyssey-3-generated environment. The company says the goal is for agents to learn from the outcomes of their own actions. The public preview covers interactive environment generation only. Robotics and autonomous-systems use needs further work.

Context

Odyssey was founded in 2023. In June 2026 it raised $310 million from investors including Amazon and AMD Ventures. Competing efforts include Google DeepMind's Genie 3 and World Labs. AMD announced in late September that it plans to acquire World Labs for roughly $8.2 billion.

What this means

The free preview puts a real-time, text-to-interactive-world model in users' hands and gives outsiders a way to test Odyssey's claims. The benchmark story is weaker than the headline suggests. The 63.37 four-run average is the figure that meets Physics-IQ's methodology, and the WorldMark rankings are self-reported. The more consequential bet is the controller strategy: one world model serving robotics, drones and games, so that simulated experience can train agents. The evidence so far is a set of company demos, with no independent replication and no pricing or API terms. Until those exist, Odyssey-3 is best treated as a credible research preview rather than a validated robotics platform.

Related Articles

model release

Reka AI releases Rho-1, a 19B-parameter omni-model for text, image, video and robot control

Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions in a single network. Reka says it uses no tool calls or external models. Context window, pricing, and benchmark scores have not been disclosed.

model release

Microsoft releases Decision-1, a Qwen3.5-9B-based model for classification and routing, at $0.042 per 1M input tokens

Microsoft has released Decision-1, a decision model built on Qwen3.5-9B for classification, evaluation, and routing. Microsoft claims 83.5% accuracy across 36 benchmarks and 85 ms latency. Input tokens cost $0.042 per million, and output tokens are free.

model release

Microsoft releases FrogNano-4B, an Apache 2.0 coding agent trained with RL on 1,500 synthetic tasks

Microsoft has released FrogNano-4B-2609, a repository-level coding agent derived from Qwen3.5-4B and published under Apache 2.0 with open weights. Microsoft says it was post-trained only with reinforcement learning on about 1,500 synthetic software-engineering tasks, with no stronger-model trajectories. It is evaluated at roughly 131K tokens of context.

model release

Qwen releases Qwen-Image-2.1-Turbo: 8-step text-to-image and editing checkpoint on a 7B architecture

Qwen has published Qwen-Image-2.1-Turbo on Hugging Face, an accelerated checkpoint of Qwen-Image-2.1 that runs text-to-image generation and image editing in 8 denoising steps. It keeps the same 7B visual generation architecture and loads through a new QwenImage21Pipeline in Diffusers.

Comments

Loading...