全量来源索引(按发布年月)
去重后 208 条一手来源(按
date年份归组;去重键为原始 URL)。
2017(1 条)
- Value Prediction Network — University of Michigan, Google Brain · 2017-07 · paper · [rl-wm] — https://arxiv.org/abs/1707.03497
2018(2 条)
- World Models (Recurrent World Models Facilitate Policy Evolution) — Google Brain, IDSIA/NNAISENSE · 2018-03 · paper · [rl-wm] — https://arxiv.org/abs/1803.10122
- Learning Latent Dynamics for Planning from Pixels (PlaNet) — Google Brain / DeepMind / University of Toronto · 2018-11 · paper · [rl-wm] — https://arxiv.org/abs/1811.04551
2019(4 条)
- Model-Based Reinforcement Learning for Atari — Google Brain / University of Illinois at Urbana-Champaign / University of Warsaw / Institute of Mathematics of the Polish Academy of Sciences / deepsense.ai / Stanford University · 2019-03 · paper · [rl-wm] — https://arxiv.org/abs/1903.00374
- When to Trust Your Model: Model-Based Policy Optimization — UC Berkeley · 2019-06 · paper · [rl-wm] — https://arxiv.org/abs/1906.08253
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (MuZero) — DeepMind · 2019-11 · paper · [rl-wm] — https://arxiv.org/abs/1911.08265
- Dream to Control: Learning Behaviors by Latent Imagination (Dreamer) — Google Brain / DeepMind / University of Toronto · 2019-12 · paper · [rl-wm] — https://arxiv.org/abs/1912.01603
2020(3 条)
- Learning to Simulate Dynamic Environments with GameGAN — NVIDIA / University of Toronto / MIT / Vector Institute · 2020-05 · paper · [game-wm] — https://arxiv.org/abs/2005.12126
- Model-based Reinforcement Learning: A Survey — Leiden University (LIACS) / TU Delft (Interactive Intelligence) · 2020-06 · paper · [survey] — https://arxiv.org/abs/2006.16712
- Mastering Atari with Discrete World Models (DreamerV2) — Google Research / DeepMind / University of Toronto · 2020-10 · paper · [rl-wm] — https://arxiv.org/abs/2010.02193
2021(6 条)
- Playable Video Generation — University of Trento / Télécom Paris (IP Paris) / Snap Inc. / Fondazione Bruno Kessler · 2021-01 · paper · [game-wm] — https://arxiv.org/abs/2101.12195
- Learning and Planning in Complex Action Spaces (Sampled MuZero) — DeepMind · 2021-04 · paper · [rl-wm] — https://arxiv.org/abs/2104.06303
- Online and Offline Reinforcement Learning by Planning with a Learned Model (MuZero Unplugged) — DeepMind · 2021-04 · paper · [rl-wm] — https://arxiv.org/abs/2104.06294
- Physion: Evaluating Physical Prediction from Vision in Humans and Machines — Stanford, MIT, UCSD · 2021-06 · paper · [method] — https://arxiv.org/abs/2106.08261
- Vector Quantized Models for Planning — DeepMind · 2021-06 · paper · [rl-wm] — https://arxiv.org/abs/2106.04615
- Mastering Atari Games with Limited Data (EfficientZero) — Tsinghua University (IIIS) · UC Berkeley · 2021-11 · paper · [rl-wm] — https://arxiv.org/abs/2111.00210
2022(9 条)
- TransDreamer: Reinforcement Learning with Transformer World Models — Rutgers University / KAIST · 2022-02 · paper · [rl-wm] — https://arxiv.org/abs/2202.09481
- Temporal Difference Learning for Model Predictive Control — UC San Diego · 2022-03 · paper · [rl-wm] — https://arxiv.org/abs/2203.04955
- Planning in Stochastic Environments with a Learned Model (Stochastic MuZero) — DeepMind · 2022-04 · paper · [rl-wm] — https://openreview.net/forum?id=X6D9bAHhBQ1
- Policy Improvement by Planning with Gumbel (Gumbel MuZero) — DeepMind · 2022-04 · paper · [method] — https://openreview.net/forum?id=bERaNdoegnO
- Iso-Dream: Isolating and Leveraging Noncontrollable Visual Dynamics in World Models — Shanghai Jiao Tong University (MoE Key Lab of Artificial Intelligence, AI Institute) · 2022-05 · paper · [rl-wm] — https://arxiv.org/abs/2205.13817
- A Path Towards Autonomous Machine Intelligence — Meta FAIR / NYU · 2022-06 · paper · [survey] — https://openreview.net/forum?id=BZ5a1r-kVsf
- DayDreamer: World Models for Physical Robot Learning — UC Berkeley · 2022-06 · paper · [rl-wm] — https://arxiv.org/abs/2206.14176
- Deep Hierarchical Planning from Pixels (Director) — Google Research / UC Berkeley / University of Toronto / Covariant · 2022-06 · paper · [rl-wm] — https://arxiv.org/abs/2206.04114
- Transformers are Sample-Efficient World Models (IRIS) — University of Geneva · 2022-09 · paper · [rl-wm] — https://arxiv.org/abs/2209.00588
2023(32 条)
- Learning Universal Policies via Text-Guided Video Generation — Google DeepMind, MIT, UC Berkeley · 2023-01 · paper · [video-wm] — https://arxiv.org/abs/2302.00111
- Mastering Diverse Domains through World Models (DreamerV3) — DeepMind / University of Toronto · 2023-01 · paper · [rl-wm] — https://arxiv.org/abs/2301.04104
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture (I-JEPA) — Meta AI (FAIR) · 2023-01 · paper · [jepa] — https://arxiv.org/abs/2301.08243
- Transformer-based World Models Are Happy With 100k Interactions (TWM) — TU Dortmund University · 2023-03 · paper · [rl-wm] — https://arxiv.org/abs/2303.07109
- Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties — MIT, Stanford, UC Berkeley, MIT-IBM Watson AI Lab, UMass Amherst · 2023-06 · paper · [method] — https://arxiv.org/abs/2306.15668
- Facing Off World Model Backbones: RNNs, Transformers, and S4 — KAIST, Rutgers University · 2023-07 · paper · [rl-wm] — https://arxiv.org/abs/2307.02064
- Learning to Model the World With Language (Dynalang) — UC Berkeley · 2023-07 · paper · [rl-wm] — https://arxiv.org/abs/2308.01399
- MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features — Meta AI (FAIR) / Inria · 2023-07 · paper · [jepa] — https://arxiv.org/abs/2307.12698
- DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving — GigaAI / Tsinghua University · 2023-09 · paper · [driving-wm] — https://arxiv.org/abs/2309.09777
- GAIA-1: A Generative World Model for Autonomous Driving — Wayve · 2023-09 · tech-report · [driving-wm] — https://arxiv.org/abs/2309.17080
- HarmonyDream: Task Harmonization Inside World Models — Tsinghua University (THUML) / Huawei Noah’s Ark Lab / Tianjin University · 2023-09 · paper · [rl-wm] — https://arxiv.org/abs/2310.00344
- DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model — Baidu Inc. · 2023-10 · paper · [driving-wm] — https://arxiv.org/abs/2310.07771
- Learning Interactive Real-World Simulators (UniSim) — UC Berkeley / Google DeepMind / MIT / University of Alberta · 2023-10 · paper · [video-wm] — https://arxiv.org/abs/2310.06114
- Learning to Act from Actionless Videos through Dense Correspondences — MIT, National Taiwan University · 2023-10 · paper · [video-wm] — https://arxiv.org/abs/2310.08576
- LightZero: A Unified Benchmark for Monte Carlo Tree Search in General Sequential Decision Scenarios — Shanghai AI Laboratory (OpenDILab), SenseTime, CUHK · 2023-10 · paper · [method] — https://arxiv.org/abs/2310.08348
- MagicDrive: Street View Generation with Diverse 3D Geometry Control — CUHK / HKUST / Huawei Noah’s Ark Lab · 2023-10 · paper · [driving-wm] — https://arxiv.org/abs/2310.02601
- STORM: Efficient Stochastic Transformer based World Models for Reinforcement Learning — Beijing Institute of Technology · 2023-10 · paper · [rl-wm] — https://arxiv.org/abs/2310.09615
- TD-MPC2: Scalable, Robust World Models for Continuous Control — UC San Diego · 2023-10 · paper · [rl-wm] — https://arxiv.org/abs/2310.16828
- Video Language Planning — Google DeepMind, MIT, UC Berkeley · 2023-10 · paper · [video-wm] — https://arxiv.org/abs/2310.10625
- A-JEPA: Joint-Embedding Predictive Architecture Can Listen — Kunlun Inc. · 2023-11 · paper · [jepa] — https://arxiv.org/abs/2311.15830
- ADriver-I: A General World Model for Autonomous Driving — MEGVII Technology / Waseda University / USTC / Mach Drive · 2023-11 · paper · [driving-wm] — https://arxiv.org/abs/2311.13549
- Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications — HAOMO.AI Technology / Shanghai Jiao Tong University · 2023-11 · paper · [driving-wm] — https://arxiv.org/abs/2311.17663
- Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion — Waabi / University of Toronto · 2023-11 · paper · [driving-wm] — https://waabi.ai/research/copilot-4d
- Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving (Drive-WM) — CASIA / UCAS / CAIR-HKISI-CAS · 2023-11 · paper · [driving-wm] — https://arxiv.org/abs/2311.17918
- OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving — Tsinghua University · 2023-11 · paper · [driving-wm] — https://arxiv.org/abs/2311.16038
- Panacea: Panoramic and Controllable Video Generation for Autonomous Driving — University of Science and Technology of China / MEGVII Technology / Mach Drive · 2023-11 · paper · [driving-wm] — https://arxiv.org/abs/2311.16813
- VBench: Comprehensive Benchmark Suite for Video Generative Models — Shanghai AI Laboratory / NTU S-Lab · 2023-11 · paper · [method] — https://arxiv.org/abs/2311.17982
- General World Models: The Next Frontier in AI Research — Runway · 2023-12 · blog · [method] — https://runwayml.com/research/introducing-general-world-models
- Understanding Physical Dynamics with Counterfactual World Modeling — Stanford University, MIT · 2023-12 · paper · [method] — https://arxiv.org/abs/2312.06721
- Visual Point Cloud Forecasting enables Scalable Autonomous Driving (ViDAR) — OpenDriveLab (Shanghai AI Lab) · 2023-12 · paper · [driving-wm] — https://arxiv.org/abs/2312.17655
- WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation — Fudan University · 2023-12 · paper · [driving-wm] — https://arxiv.org/abs/2312.02934
- WonderJourney: Going from Anywhere to Everywhere — Stanford University, Google Research · 2023-12 · paper · [video-wm] — https://arxiv.org/abs/2312.03884
2024(61 条)
- WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens — GigaAI / Tsinghua University · 2024-01 · paper · [video-wm] — https://arxiv.org/abs/2401.09985
- Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning — Princeton University, UT Austin, Meta AI (FAIR) · 2024-02 · paper · [rl-wm] — https://arxiv.org/abs/2402.03570
- Improving Token-Based World Models with Parallel Observation Prediction — Technion – Israel Institute of Technology, ByteDance · 2024-02 · paper · [rl-wm] — https://arxiv.org/abs/2402.05643
- Revisiting Feature Prediction for Learning Visual Representations from Video (V-JEPA) — Meta FAIR · 2024-02 · paper · [jepa] — https://arxiv.org/abs/2404.08471
- World Model on Million-Length Video And Language With Blockwise RingAttention — UC Berkeley · 2024-02 · paper · [video-wm] — https://arxiv.org/abs/2402.08268
- 3D-VLA: A 3D Vision-Language-Action Generative World Model — UMass Amherst / MIT / MIT-IBM Watson AI Lab / UCLA / Shanghai Jiao Tong Univ. / South China Univ. of Tech. / Wuhan Univ. · 2024-03 · paper · [video-wm] — https://arxiv.org/abs/2403.09631
- DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation — GigaAI / Institute of Automation, Chinese Academy of Sciences (CASIA) · 2024-03 · paper · [driving-wm] — https://arxiv.org/abs/2403.06845
- EfficientZero V2: Mastering Discrete and Continuous Control with Limited Data — Tsinghua University (IIIS), Shanghai Qi Zhi Institute · 2024-03 · paper · [rl-wm] — https://arxiv.org/abs/2403.00564
- GenAD: Generalized Predictive Model for Autonomous Driving — OpenDriveLab (Shanghai AI Lab) · 2024-03 · paper · [driving-wm] — https://arxiv.org/abs/2403.09630
- Learning and Leveraging World Models in Visual Representation Learning — Meta FAIR · 2024-03 · paper · [jepa] — https://arxiv.org/abs/2403.00504
- Mastering Memory Tasks with World Models — Mila; Universite de Montreal; Polytechnique Montreal; Dalhousie University · 2024-03 · paper · [rl-wm] — https://arxiv.org/abs/2403.04253
- S-JEPA: towards seamless cross-dataset transfer through dynamic spatial attention — Donders Institute (Radboud University) / Inria · 2024-03 · paper · [jepa] — https://arxiv.org/abs/2403.11772
- Scaling Instructable Agents Across Many Simulated Worlds — Google DeepMind · 2024-03 · tech-report · [platform] — https://arxiv.org/abs/2404.10179
- Point-JEPA: A Joint Embedding Predictive Architecture for Self-Supervised Learning on Point Cloud — Graphics and Spatial Computing Lab, Saint Mary’s University (Canada) · 2024-04 · paper · [jepa] — https://arxiv.org/abs/2404.16432
- ReZero: Boosting MCTS-based Algorithms by Backward-view and Entire-buffer Reanalyze — Shanghai Artificial Intelligence Laboratory / Xi’an Jiaotong University / SenseTime Research · 2024-04 · paper · [method] — https://arxiv.org/abs/2404.16364
- RoboDreamer: Learning Compositional World Models for Robot Imagination — HKUST, MIT, UCSD, UCF, UMass Amherst, MIT-IBM Watson AI Lab · 2024-04 · paper · [video-wm] — https://arxiv.org/abs/2404.12377
- Diffusion for World Modeling: Visual Details Matter in Atari (DIAMOND) — University of Geneva / University of Edinburgh / Microsoft Research · 2024-05 · paper · [rl-wm] — https://arxiv.org/abs/2405.12399
- DriveWorld: 4D Pre-trained Scene Understanding via World Models for Autonomous Driving — Defense Innovation Institute / Peking University / HKUST et al. · 2024-05 · paper · [driving-wm] — https://arxiv.org/abs/2405.04390
- MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes — CUHK / HKUST / Huawei Noah’s Ark Lab · 2024-05 · paper · [driving-wm] — https://arxiv.org/abs/2405.14475
- OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving — Beihang University / Tsinghua University · 2024-05 · paper · [driving-wm] — https://arxiv.org/abs/2405.20337
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability — OpenDriveLab (Shanghai AI Lab) / HKUST / U. Tübingen · 2024-05 · paper · [driving-wm] — https://arxiv.org/abs/2405.17398
- iVideoGPT: Interactive VideoGPTs are Scalable World Models — Tsinghua University · 2024-05 · paper · [video-wm] — https://arxiv.org/abs/2405.15223
- Efficient World Models with Context-Aware Tokenization (Δ-IRIS) — University of Geneva · 2024-06 · paper · [rl-wm] — https://arxiv.org/abs/2406.19320
- IRASim: A Fine-Grained World Model for Robot Manipulation — ByteDance Research · 2024-06 · paper · [video-wm] — https://arxiv.org/abs/2406.14540
- OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation — Fudan University, ByteDance Inc. · 2024-06 · paper · [method] — https://arxiv.org/abs/2406.09399
- Pandora: Towards General World Model with Natural Language Actions and Video States — Maitrix.org / UC San Diego / MBZUAI / Carnegie Mellon University · 2024-06 · paper · [video-wm] — https://arxiv.org/abs/2406.09455
- UniZero: Generalized and Efficient Planning with Scalable Latent World Models — Shanghai Artificial Intelligence Laboratory / SenseTime Research / CUHK / Shanghai Jiao Tong University · 2024-06 · paper · [rl-wm] — https://arxiv.org/abs/2406.10667
- Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation — Westlake University / Li Auto Inc. · 2024-06 · paper · [driving-wm] — https://arxiv.org/abs/2406.01349
- VideoPhy: Evaluating Physical Commonsense for Video Generation — UCLA, Google Research · 2024-06 · paper · [method] — https://arxiv.org/abs/2406.03520
- WonderWorld: Interactive 3D Scene Generation from a Single Image — Stanford University, MIT · 2024-06 · paper · [video-wm] — https://arxiv.org/abs/2406.09394
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion — MIT · 2024-07 · paper · [method] — https://arxiv.org/abs/2407.01392
- PWM: Policy Learning with Multi-Task World Models — Georgia Institute of Technology, UC San Diego · 2024-07 · paper · [rl-wm] — https://arxiv.org/abs/2407.02466
- Diffusion Models Are Real-Time Game Engines (GameNGen) — Google Research / Google DeepMind / Tel Aviv University · 2024-08 · paper · [game-wm] — https://arxiv.org/abs/2408.14837
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving — Zhejiang University / Huawei Technologies · 2024-08 · paper · [driving-wm] — https://arxiv.org/abs/2408.14197
- Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation — Sony Computer Science Laboratories – Paris (Sony CSL Paris) · 2024-08 · paper · [jepa] — https://arxiv.org/abs/2408.02514
- 1X World Model — 1X Technologies · 2024-09 · blog · [rl-wm] — https://www.1x.tech/discover/1x-world-model
- Brain-JEPA: Brain Dynamics Foundation Model with Gradient Positioning and Spatiotemporal Masking — National University of Singapore · 2024-09 · paper · [jepa] — https://arxiv.org/abs/2409.19407
- AVID: Adapting Video Diffusion Models to World Models — Imperial College London, Microsoft Research · 2024-10 · paper · [method] — https://arxiv.org/abs/2410.12822
- DOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model — Horizon Robotics / CASIA / HKUST · 2024-10 · paper · [driving-wm] — https://arxiv.org/abs/2410.10429
- DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation — GigaAI / Institute of Automation, Chinese Academy of Sciences · 2024-10 · paper · [driving-wm] — https://arxiv.org/abs/2410.13571
- LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior — University of Maryland, College Park · 2024-10 · paper · [method] — https://arxiv.org/abs/2410.21264
- Oasis: A Universe in a Transformer — Decart · 2024-10 · blog · [game-wm] — https://oasis-model.github.io/
- T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular Data — Université Paris-Saclay, CNRS, CentraleSupélec (LISN); Emobot · 2024-10 · paper · [jepa] — https://arxiv.org/abs/2410.05016
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation — Shanghai AI Laboratory (OpenGVLab), Shanghai Jiao Tong University, HKU, CUHK · 2024-10 · paper · [method] — https://arxiv.org/abs/2410.05363
- WorldSimBench: Towards Video Generation Models as World Simulators — Shanghai AI Laboratory, CUHK-Shenzhen, Beihang University, HKU · 2024-10 · paper · [method] — https://arxiv.org/abs/2410.18072
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning — New York University / Meta FAIR · 2024-11 · paper · [jepa] — https://arxiv.org/abs/2411.04983
- GameGen-X: Interactive Open-world Game Video Generation — The Hong Kong University of Science and Technology / University of Science and Technology of China / The Chinese University of Hong Kong · 2024-11 · paper · [game-wm] — https://arxiv.org/abs/2411.00769
- MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control — CUHK / HKUST / Huawei Cloud / Huawei Noah’s Ark Lab · 2024-11 · paper · [driving-wm] — https://arxiv.org/abs/2411.13807
- ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration — GigaAI / Peking University / Li Auto Inc. / CASIA · 2024-11 · paper · [driving-wm] — https://arxiv.org/abs/2411.19548
- Understanding World or Predicting Future? A Comprehensive Survey of World Models — Tsinghua University (FIB Lab, Jingtao Ding / Yong Li) · 2024-11 · paper · [survey] — https://arxiv.org/abs/2411.14499
- Doe-1: Closed-Loop Autonomous Driving with Large World Model — Tsinghua University · 2024-12 · paper · [driving-wm] — https://arxiv.org/abs/2412.09627
- DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT — HKUST / Horizon Robotics · 2024-12 · paper · [driving-wm] — https://arxiv.org/abs/2412.19505
- GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction — Tsinghua University · 2024-12 · paper · [driving-wm] — https://arxiv.org/abs/2412.10373
- GenEx: Generating an Explorable World — Johns Hopkins University · 2024-12 · paper · [game-wm] — https://arxiv.org/abs/2412.09624
- Genesis: A Generative and Universal Physics Engine for Robotics and Beyond — Genesis Authors(CMU / Stanford / MIT / NVIDIA 等近 20 家机构学术协作,2024-12 后团队创立 Genesis AI 接手) · 2024-12 · model-card · [platform] — https://github.com/Genesis-Embodied-AI/Genesis
- Genie 2: A Large-Scale Foundation World Model — Google DeepMind · 2024-12 · blog · [game-wm] — https://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model/
- InfinityDrive: Breaking Time Limits in Driving World Models — SenseAuto Research / Tsinghua University · 2024-12 · paper · [driving-wm] — https://arxiv.org/abs/2412.01522
- Navigation World Models — FAIR at Meta / New York University / UC Berkeley · 2024-12 · paper · [video-wm] — https://arxiv.org/abs/2412.03572
- Owl-1: Omni World Model for Consistent Long Video Generation — Tsinghua University / Kuaishou Technology · 2024-12 · paper · [video-wm] — https://arxiv.org/abs/2412.09600
- Playable Game Generation — Tencent · 2024-12 · paper · [game-wm] — https://arxiv.org/abs/2412.00887
- The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control — Alibaba Group (Tongyi Lab) / University of Hong Kong / University of Waterloo / Vector Institute · 2024-12 · paper · [game-wm] — https://arxiv.org/abs/2412.03568
2025(72 条)
- Do generative video models understand physical principles? — Google DeepMind · 2025-01 · paper · [method] — https://arxiv.org/abs/2501.09038
- EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation — AgiBot (Zhiyuan Robotics), Shanghai AI Laboratory · 2025-01 · paper · [video-wm] — https://arxiv.org/abs/2501.01895
- GameFactory: Creating New Games with Generative Interactive Videos — The University of Hong Kong / Kuaishou Technology · 2025-01 · paper · [game-wm] — https://arxiv.org/abs/2501.08325
- HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation — Huazhong University of Science & Technology / MEGVII Technology / Mach Drive / The University of Hong Kong · 2025-01 · paper · [driving-wm] — https://arxiv.org/abs/2501.14729
- PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding — University of Southern California, UC Berkeley, Toyota Research Institute · 2025-01 · paper · [method] — https://arxiv.org/abs/2501.16411
- Ray2 (Luma Dream Machine video model) — Luma AI · 2025-01 · blog · [video-wm] — https://lumalabs.ai/ray2
- VideoWorld: Exploring Knowledge Learning from Unlabeled Videos — Beijing Jiaotong University · University of Science and Technology of China · ByteDance Seed · 2025-01 · paper · [video-wm] — https://arxiv.org/abs/2501.09781
- History-Guided Video Diffusion — MIT · 2025-02 · paper · [method] — https://arxiv.org/abs/2502.06764
- Intuitive physics understanding emerges from self-supervised pretraining on natural videos — Meta FAIR · 2025-02 · paper · [jepa] — https://arxiv.org/abs/2502.11831
- Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models — NYU / Meta FAIR / Brown / Toronto / Genentech · 2025-02 · paper · [jepa] — https://arxiv.org/abs/2502.14819
- Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving — Institute for AI Industry Research (AIR), Tsinghua University · 2025-02 · paper · [driving-wm] — https://arxiv.org/abs/2502.07309
- The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey — Huazhong University of Science and Technology / Baidu Inc. · 2025-02 · paper · [survey] — https://arxiv.org/abs/2502.10498
- Muse) — Microsoft Research · 2025-02 · paper · [game-wm] — https://www.nature.com/articles/s41586-025-08600-3
- WorldModelBench: Judging Video Generation Models As World Models — UC Berkeley, UC San Diego, NVIDIA, MIT · 2025-02 · paper · [survey] — https://arxiv.org/abs/2502.20694
- AdaWorld: Learning Adaptable World Models with Latent Actions — HKUST, Harvard University, UMass Amherst, MIT-IBM Watson AI Lab · 2025-03 · paper · [rl-wm] — https://arxiv.org/abs/2503.18938
- Aether: Geometric-Aware Unified World Modeling — Shanghai AI Laboratory · 2025-03 · paper · [video-wm] — https://arxiv.org/abs/2503.18945
- Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning — NVIDIA · 2025-03 · paper · [method] — https://arxiv.org/abs/2503.15558
- Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control — NVIDIA · 2025-03 · tech-report · [platform] — https://arxiv.org/abs/2503.14492
- GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving — Wayve · 2025-03 · tech-report · [driving-wm] — https://arxiv.org/abs/2503.20523
- Impossible Videos — Show Lab, National University of Singapore · 2025-03 · paper · [method] — https://arxiv.org/abs/2503.14378
- Learning Transformer-based World Models with Contrastive Predictive Coding (TWISTER) — University of Würzburg (Computer Vision Lab, CAIDAS & IFI) · 2025-03 · paper · [rl-wm] — https://arxiv.org/abs/2503.04416
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness — Shanghai AI Laboratory / NTU S-Lab · 2025-03 · paper · [method] — https://arxiv.org/abs/2503.21755
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation — UCLA, Google Research · 2025-03 · paper · [method] — https://arxiv.org/abs/2503.06800
- WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation — 360 AI Research (Qihoo 360), Sun Yat-Sen University (Shenzhen Campus), Peng Cheng Laboratory · 2025-03 · paper · [method] — https://arxiv.org/abs/2503.08153
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft — Microsoft Research · 2025-04 · paper · [game-wm] — https://arxiv.org/abs/2504.08388
- TesserAct: Learning 4D Embodied World Models — UMass Amherst, HKUST, Harvard University · 2025-04 · paper · [video-wm] — https://arxiv.org/abs/2504.20995
- WorldMem: Long-term Consistent World Simulation with Memory — Nanyang Technological University (S-Lab) / Peking University / Shanghai AI Laboratory · 2025-04 · paper · [game-wm] — https://arxiv.org/abs/2504.12369
- WorldScore: A Unified Evaluation Benchmark for World Generation — Stanford University (SVL) · 2025-04 · paper · [survey] — https://arxiv.org/abs/2504.00983
- Odyssey-1: A Playable World Model — Odyssey · 2025-05 · blog · [video-wm] — https://odyssey.ml/introducing-odyssey-1
- Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models — NVIDIA · 2025-06 · tech-report · [driving-wm] — https://arxiv.org/abs/2506.09042
- Cosmos-Predict2: A Suite of Diffusion-based World Foundation Models Available in 2B, and 14B — NVIDIA · 2025-06 · model-card · [platform] — https://github.com/nvidia-cosmos/cosmos-predict2
- Epona: Autoregressive Diffusion World Model for Autonomous Driving — Horizon Robotics / Tsinghua University · 2025-06 · paper · [driving-wm] — https://arxiv.org/abs/2506.24113
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition — Tencent Hunyuan · 2025-06 · paper · [game-wm] — https://arxiv.org/abs/2506.17201
- Matrix-Game: Interactive World Foundation Model — Skywork AI · 2025-06 · tech-report · [game-wm] — https://arxiv.org/abs/2506.18701
- PlayerOne: Egocentric World Simulator — HKU / Alibaba DAMO Academy / Hupan Lab / HUST · 2025-06 · paper · [game-wm] — https://arxiv.org/abs/2506.09995
- RoboScape: Physics-informed Embodied World Model — Tsinghua University, Manifold AI · 2025-06 · paper · [video-wm] — https://arxiv.org/abs/2506.23135
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Meta FAIR · 2025-06 · paper · [jepa] — https://arxiv.org/abs/2506.09985
- Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation — Tencent Hunyuan · 2025-06 · paper · [platform] — https://arxiv.org/abs/2506.04225
- HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels — Tencent Hunyuan · 2025-07 · paper · [platform] — https://arxiv.org/abs/2507.21809
- I2-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting — Xi’an Jiaotong University · 2025-07 · paper · [method] — https://arxiv.org/abs/2507.09144
- Mirage: Real-Time AI-Native UGC Game Engine — Dynamics Lab · 2025-07 · blog · [game-wm] — https://blog.dynamicslab.ai/
- MirageLSD: The First Live-Stream Diffusion AI Video Model — Decart · 2025-07 · blog · [video-wm] — https://decart.ai/publications/mirage
- MirageLSD: Zero-Latency, Real-Time, Infinite Video Generation — Decart · 2025-07 · blog · [video-wm] — https://mirage.decart.ai/
- PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models — University of Maryland, NVIDIA · 2025-07 · paper · [method] — https://arxiv.org/abs/2507.13428
- Vidar: Embodied Video Diffusion Model for Generalist Manipulation — Tsinghua University, ShengShu Technology · 2025-07 · paper · [video-wm] — https://arxiv.org/abs/2507.12898
- World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model — CASIA / Li Auto / PCL / NUS / Tsinghua · 2025-07 · paper · [driving-wm] — https://arxiv.org/abs/2507.00603
- Yume: An Interactive World Generation Model — Shanghai AI Laboratory / Fudan University / Shanghai Innovation Institute · 2025-07 · paper · [game-wm] — https://arxiv.org/abs/2507.17744
- Genie 3: A new frontier for world models — Google DeepMind · 2025-08 · blog · [platform] — https://deepmind.google/discover/blog/genie-3-a-new-frontier-for-world-models/
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation — AgiBot (Zhiyuan Robotics) · 2025-08 · tech-report · [platform] — https://arxiv.org/abs/2508.05635
- Matrix-3D: Omnidirectional Explorable 3D World Generation — Skywork AI · 2025-08 · paper · [game-wm] — https://arxiv.org/abs/2508.08086
- Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World Model — Skywork AI · 2025-08 · tech-report · [game-wm] — https://arxiv.org/abs/2508.13009
- Yan: Foundational Interactive Video Generation — Tencent · 2025-08 · paper · [game-wm] — https://arxiv.org/abs/2508.08601
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures — Brown University / NYU / Atlassian · 2025-09 · paper · [jepa] — https://arxiv.org/abs/2509.14252
- Ray3 and Luma’s Multimodal World-Model Program — Luma AI · 2025-09 · blog · [video-wm] — https://lumalabs.ai/news/ray3
- WoW: Towards a World omniscient World model Through Embodied Interaction — Beijing Innovation Center of Humanoid Robotics (X-Humanoid) · Peking University · HKUST · 2025-09 · paper · [video-wm] — https://arxiv.org/abs/2509.22642
- WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving — Nankai University / Xiaomi EV / Nanjing University, Suzhou · 2025-09 · paper · [driving-wm] — https://arxiv.org/abs/2509.23402
- Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models — Waymo / Google DeepMind · 2025-10 · paper · [method] — https://arxiv.org/abs/2510.06209
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction — Dalian University of Technology · 2025-10 · paper · [driving-wm] — https://arxiv.org/abs/2510.19654
- JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation — Jiangsu University 领衔(跨机构:USC / 北大 / 南洋理工 / Brown / UCSD / 石溪大学 / 中科院大学) · 2025-10 · paper · [method] — https://arxiv.org/abs/2510.00974
- RTFM: A Real-Time Frame Model — World Labs · 2025-10 · blog · [platform] — https://www.worldlabs.ai/blog/rtfm
- Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks — Peking University / Xiaomi EV / Huazhong University of Science and Technology · 2025-10 · paper · [driving-wm] — https://arxiv.org/abs/2510.19195
- SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries — Huazhong University of Science and Technology / AIR, Tsinghua University / Lenovo Group · 2025-10 · paper · [driving-wm] — https://arxiv.org/abs/2510.17482
- World Simulation with Video Foundation Models for Physical AI (Cosmos-Predict2.5 & Cosmos-Transfer2.5) — NVIDIA · 2025-10 · tech-report · [platform] — https://arxiv.org/abs/2511.00062
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics — Brown University / Meta FAIR · 2025-11 · paper · [jepa] — https://arxiv.org/abs/2511.08544
- Marble: A Multimodal World Model — World Labs · 2025-11 · blog · [platform] — https://www.worldlabs.ai/blog/marble-world-model
- PAN: A World Model for General, Interactable, and Long-Horizon World Simulation — MBZUAI Institute of Foundation Models (IFM) · 2025-11 · tech-report · [method] — https://arxiv.org/abs/2511.09057
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds — Google DeepMind · 2025-11 · tech-report · [rl-wm] — https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/
- SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model — Tongji University / Li Auto Inc · 2025-11 · paper · [driving-wm] — https://arxiv.org/abs/2511.22039
- GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation — Shanghai Jiao Tong University / Tsinghua University / MEGVII Technology / Mach Drive · 2025-12 · paper · [driving-wm] — https://arxiv.org/abs/2512.23180
- GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation — The University of Hong Kong / Huawei Noah’s Ark Lab / Huazhong University of Science and Technology · 2025-12 · paper · [driving-wm] — https://arxiv.org/abs/2512.12751
- JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention — Academic collaboration (CMU / Univ. of Bristol / Stanford / NYU / Northeastern / James Silberrad Brown Center for AI, incl. Yann LeCun) · 2025-12 · paper · [method] — https://arxiv.org/abs/2512.07168
- JEPA-Reasoner: Decoupling Latent Reasoning from Token Generation — unknown (未披露机构,作者含 David P. Woodruff) · 2025-12 · paper · [jepa] — https://arxiv.org/abs/2512.19171
2026(18 条)
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning — NVIDIA · 2026-01 · paper · [rl-wm] — https://research.nvidia.com/labs/dir/cosmos-policy/
- DriveWorld-VLA: Unified Latent-Space World Modeling with Vision–Language–Action for Autonomous Driving — China (academic; institutional affiliation not disclosed in paper) · 2026-02 · paper · [driving-wm] — https://arxiv.org/abs/2602.06521
- JEPA-VLA: Video Predictive Embedding is Needed for VLA Models — academic (unverified) · 2026-02 · paper · [jepa] — https://arxiv.org/abs/2602.11832
- Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation (WorldDrive) — University of Macau (SKL-IOTSC) / Afari Intelligent Drive · 2026-03 · paper · [driving-wm] — https://arxiv.org/abs/2603.14948
- Hermes++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation — Huazhong University of Science and Technology / Mach Drive / The University of Hong Kong · 2026-04 · paper · [driving-wm] — https://arxiv.org/abs/2604.28196
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory — Skywork AI · 2026-04 · tech-report · [game-wm] — https://arxiv.org/abs/2604.08995
- MultiWorld: Scalable Multi-Agent Multi-View Video World Models — The University of Hong Kong; Sreal AI · 2026-04 · paper · [video-wm] — https://arxiv.org/abs/2604.18564
- When Does LeJEPA Learn a World Model? — Cold Spring Harbor Laboratory / New York University / Brown University · 2026-05 · paper · [survey] — https://arxiv.org/abs/2605.26379
- Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving — Xiaomi · 2026-05 · tech-report · [driving-wm] — https://arxiv.org/abs/2605.18137
- DreamX-World 1.0: A General-Purpose Interactive World Model — AMAP-ML / DreamX Team(高德地图,阿里巴巴集团) · 2026-06 · tech-report · [game-wm] — https://arxiv.org/abs/2606.16993
- NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation — NVIDIA · 2026-06 · tech-report · [driving-wm] — https://arxiv.org/abs/2606.03159
- World Action Models: A Survey — National University of Singapore · 2026-06 · paper · [survey] — https://arxiv.org/abs/2606.20781
- ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space — AMAP CV Lab, Alibaba Group · 2026-07 · tech-report · [platform] — https://arxiv.org/abs/2607.11673
- AlayaWorld: Long-Horizon and Playable Video World Generation — Alaya Lab (Shanda AI Research Tokyo) · 2026-07 · tech-report · [game-wm] — https://arxiv.org/abs/2607.06291
- GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation — GigaAI · 2026-07 · paper · [platform] — https://arxiv.org/abs/2607.02642
- MoWorld: A Flash World Model — Moxin Technology / KOKONI 3D(联合浙江大学) · 2026-07 · tech-report · [video-wm] — https://arxiv.org/abs/2607.06216
- RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation — Alibaba DAMO Academy · 2026-07 · paper · [rl-wm] — https://arxiv.org/abs/2607.06559
- RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation — Alibaba DAMO Academy · 2026-07 · paper · [rl-wm] — https://arxiv.org/abs/2607.06558