Literature
Papers
The papers behind the world models in the catalog — each linked to the model it introduced, the data it used, and what it was evaluated on.
Cosmos World Foundation Model Platform for Physical AI
NVIDIA
Platform of pretrained diffusion and autoregressive video world models, tokenizers, and a curation pipeline for physical AI.
1 relationships
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Meta AI (FAIR)
Action-free joint-embedding predictive video model pretrained on 1M+ hours of video, post-trained for robot planning.
3 relationships
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
Wayve
Latent diffusion world model generating controllable, multi-camera driving video from structured inputs.
1 relationships
Mastering Diverse Domains through World Models
Google DeepMind
DreamerV3 — one model-based RL configuration that masters 150+ tasks and mines diamonds in Minecraft from scratch.
2 relationships
Genie: Generative Interactive Environments
Google DeepMind
The first foundation world model trained unsupervised from internet video, generating action-controllable worlds.
2 relationships
Diffusion Models Are Real-Time Game Engines (GameNGen)
A neural game engine that runs DOOM interactively at ~20 FPS, with the next frame generated by a diffusion model conditioned on past frames and actions.
1 relationships
Diffusion for World Modeling: Visual Details Matter in Atari (DIAMOND)
University of Geneva / Edinburgh / Microsoft Research
Trains RL agents inside a diffusion world model, showing that high-fidelity visual detail materially improves downstream policy performance on Atari.
1 relationships
Navigation World Models
Meta / New York University
A controllable video world model for navigation that imagines future observations from planned trajectories and can be used for planning.
1 relationships
iVideoGPT: Interactive VideoGPTs are Scalable World Models
Tsinghua University
A scalable autoregressive transformer world model with a compressive tokenizer, pretrained on large video corpora for interactive prediction and control.
1 relationships