Research Note
World Models and AI Video Research Note
This note supports a technical explainer that separates observed output behavior, vendor language, research definitions, interactive simulation, and hidden architecture.
World Models and AI Video Research Note
Editorial use
This note supports a technical explainer that separates observed output behavior, vendor language, research definitions, interactive simulation, and hidden architecture.
Evidence record
| Source | Relevant claim | Boundary |
|---|---|---|
| Ha and Schmidhuber, World Models | A learned compressed spatial and temporal representation of an environment can support control | Specific reinforcement-learning architecture, not a definition binding every later use |
| Google DeepMind Genie | An action-controllable generated environment trained from video is closer to an interactive world model | Different objective and interface from ordinary text-to-video |
| Google DeepMind Genie 2 | Vendor describes action-controllable, playable 3D environments for agent training | Research demonstration and vendor framing |
| OpenAI Sora technical report | OpenAI presented qualitative simulation capabilities and limitations | Does not disclose a literal physics-engine module |
| Sora 2 launch | OpenAI claimed improved physical accuracy and world-simulation capability | Product and aspiration language |
| Sora 2 System Card | OpenAI called Sora 2 a step toward models that simulate physical-world complexity more accurately | Does not establish general-purpose simulation or action control |
| Veo 3.1 | Google claims improved physics, consistency, control, and prompt adherence | Output capability claim, not a disclosed world-model architecture |
Conceptual ladder
Physical plausibility means an output looks believable. Temporal coherence means objects and attributes persist across frames. Learned dynamics means a model predicts how a representation changes. Counterfactual control means changing an action or condition produces a stable, testable consequence. Interactive simulation supports ongoing action-conditioned evolution. A general world model would require broader, reliable representation and prediction across environments and interventions.
Success at one rung is not proof of the next.
Correction to the episode
The raw transcript says Sora 2 had a built-in physics engine and legitimately understood the external world. The official record supports narrower language: OpenAI claimed more accurate physics, improved real-world dynamics, and progress toward world simulation. It did not document a separate physics-engine component or general understanding.
Publication boundary
The public explainer can accurately map terms and source language. It should remain editorial-review until a qualified technical reviewer checks the definitions, architecture boundary, and counterfactual claims.
Sources
Follow the evidence.
- deepmind.google: veodeepmind.google
- deepmind.google: veo 3 1 litedeepmind.google
- deepmind.google: model cardsdeepmind.google
- openai.com: sora 2 system cardopenai.com
- deploymentsafety.openai.com: overview of sora 2deploymentsafety.openai.com
- uspto.gov: copyright and ai digital replicas report part oneuspto.gov
- copyright.gov: Copyright and Artificial Intelligence Part 2 Copyrightability Reportcopyright.gov
- openai.com: creating with sora safelyopenai.com
- copyright.gov: aicopyright.gov
- openai.com: sora 2openai.com
- uspto.gov: name image and likenessuspto.gov