AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
Home
❯
world model
❯
2025
文件夹: world-model/2025
此文件夹下有72条笔记。
2026年7月16日
AdaWorld: Learning Adaptable World Models with Latent Actions
world-model
latent-action
unsupervised-action-learning
action-transfer
few-shot-adaptation
stable-video-diffusion
model-predictive-control
beta-vae
icml-2025
2026年7月16日
Aether: Geometric-Aware Unified World Modeling
world-model
4d-reconstruction
video-diffusion
action-conditioned-prediction
visual-planning
synthetic-data
zero-shot-transfer
camera-pose
cogvideox
2026年7月16日
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
world-model
autonomous-driving
synthetic-data
controlnet
multi-view
lidar-generation
diffusion
dit
cosmos
data-flywheel
2026年7月16日
World Simulation with Video Foundation Models for Physical AI (Cosmos-Predict2.5 & Cosmos-Transfer2.5)
world-model
physical-ai
video-foundation-model
flow-matching
controlnet
robotics
autonomous-driving
sim2real
cosmos
2026年7月16日
Cosmos-Predict2: A Suite of Diffusion-based World Foundation Models Available in 2B, and 14B
world-model
physical-ai
video-generation
text-to-image
diffusion-transformer
robotics
autonomous-driving
natten
action-conditioned
cosmos
2026年7月16日
MirageLSD: The First Live-Stream Diffusion AI Video Model
world-model
live-stream-diffusion
real-time-video
diffusion-forcing
autoregressive
zero-latency
video-transformation
mega-kernel
shortcut-distillation
2026年7月16日
MirageLSD: Zero-Latency, Real-Time, Infinite Video Generation
world-model
video-to-video
real-time-generation
diffusion-forcing
live-stream-diffusion
autoregressive-video
world-transformation
streaming
decart
mirage
2026年7月16日
SIMA 2: A Generalist Embodied Agent for Virtual Worlds
embodied-agent
vla
generalist-agent
gemini
self-improvement
world-model
genie-3
rlvr
keyboard-mouse
3d-games
2026年7月16日
Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
world-model
autonomous-driving
video-generation
end-to-end-planning
co-evaluation
behavior-permutation-test
diffusion-transformer
vlm-planner
synthetic-data-augmentation
ood-generalization
2026年7月16日
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
world-model
video-diffusion
robotic-manipulation
multi-view
free-anchor-view
4d-gaussian-splatting
sim2real
chunk-autoregressive
sparse-memory
libero
2026年7月16日
Epona: Autoregressive Diffusion World Model for Autonomous Driving
driving-world-model
autoregressive-diffusion
diffusion-transformer
rectified-flow
trajectory-planning
long-horizon-generation
chain-of-forward
deep-compression-autoencoder
nuplan
nuscenes
2026年7月16日
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
world-model
autonomous-driving
latent-diffusion
flow-matching
multi-camera
video-tokenizer
controllable-generation
safety-critical
synthetic-data
2026年7月16日
GameFactory: Creating New Games with Generative Interactive Videos
world-model
game-generation
minecraft
video-diffusion
action-control
scene-generalization
lora
autoregressive-generation
diffusion-forcing
dataset
2026年7月16日
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
world-model
autonomous-driving
3d-gaussian-splatting
vision-language-model
visual-grounding
scene-understanding
video-generation
langsplat
dual-condition-diffusion
cvpr2026
2026年7月16日
Genie 3: A new frontier for world models
world-model
interactive
real-time
autoregressive
video-generation
embodied-agent
agi
promptable-world-events
sima
2026年7月16日
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
world-model
robotic-manipulation
video-diffusion
vla
flow-matching
dual-arm
cross-embodiment
neural-simulator
benchmark
2026年7月16日
GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
driving-world-model
4d-occupancy
tri-plane-vae
occupancy-forecasting
multi-view-video-generation
wan2.1
diffusion-transformer
mutual-control-attention
nuscenes
cvpr2026
2026年7月16日
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
driving-world-model
bev-representation
world-queries
internvl2
point-cloud-forecasting
scene-understanding
unified-understanding-generation
nuscenes
iccv2025
2026年7月16日
History-Guided Video Diffusion
world-model
video-diffusion
diffusion-forcing
history-guidance
classifier-free-guidance
dit
long-video-generation
autoregressive-rollout
icml-2025
2026年7月16日
HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
world-model
3d-generation
panorama
diffusion-transformer
mesh-export
scene-layering
gaussian-splatting
text-to-3d
image-to-3d
open-source
2026年7月16日
Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
world-model
video-diffusion
rgb-d-generation
camera-control
3d-reconstruction
autoregressive-generation
point-cloud
scene-exploration
hunyuanvideo
2026年7月16日
I2-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting
world-model
autonomous-driving
4d-occupancy
tokenization
residual-quantization
occupancy-forecasting
nuscenes
occ3d
iccv2025
2026年7月16日
Impossible Videos
world-model
video-generation
video-understanding
benchmark
counterfactual
physical-commonsense
Video-LLM
taxonomy
ICML2025
2026年7月16日
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
jepa
world-model
intuitive-physics
violation-of-expectation
self-supervised
latent-prediction
video-representation
multimodal-llm-baseline
lecun
2026年7月16日
JEPA-Reasoner: Decoupling Latent Reasoning from Token Generation
jepa
latent-reasoning
coconut
gsm8k
decoupled-architecture
talker-model
error-containment
ema-target-encoder
cosine-similarity-loss
2026年7月16日
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
jepa
speech-tokenizer
self-supervised-learning
finite-scalar-quantization
density-adaptive-attention
hifi-gan
conformer
neural-audio-codec
representation-learning
2026年7月16日
JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
jepa
text-to-image
cross-attention
flow-matching
masked-prediction
imagenet
multimodal-fusion
clip
mar
non-generative-lineage
2026年7月16日
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
jepa
ssl
self-supervised
isotropic-gaussian
sigreg
epps-pulley
lecun
balestriero
collapse-free
2026年7月16日
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
jepa
llm-finetuning
llm-pretraining
embedding-space-objective
next-token-prediction
contrastive-representation-learning
custom-attention-mask
loss-dropout
lora
yann-lecun
2026年7月16日
Ray2 (Luma Dream Machine video model)
video-generation
t2v
i2v
closed-source
dream-machine
luma-ai
ray2
camera-motion
keyframes
world-model-narrative
2026年7月16日
Ray3 and Luma's Multimodal World-Model Program
world-model
video-generation
reasoning-video-model
hdr-video
luma-ai
ray3
dream-machine
multimodal-agi
2026年7月16日
Matrix-3D: Omnidirectional Explorable 3D World Generation
world-model
panoramic-video
3d-gaussian-splatting
scene-generation
diffusion-transformer
video-diffusion
feed-forward-reconstruction
synthetic-dataset
2026年7月16日
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
world-model
generative-latent-prediction
glp
autoregressive
video-diffusion
causal-swin-dpm
qwen2.5-vl
wan2.1
long-horizon
simulative-reasoning
2026年7月16日
World and Human Action Models towards gameplay ideation (WHAM / Muse)
world-model
game-wm
autoregressive
transformer
vqgan
tokenizer
action-conditioning
bleeding-edge
nature
muse
2026年7月16日
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
world-model
minecraft
autoregressive-transformer
vq-vae
action-conditioning
parallel-decoding
diagonal-decoding
real-time-interaction
vpt-dataset
controllability-metric
2026年7月16日
Mirage: Real-Time AI-Native UGC Game Engine
world-model
game-wm
ugc
autoregressive-diffusion
causal-transformer
distillation
kv-cache
cloud-gaming
real-time-generation
startup
2026年7月16日
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
physical-ai
vlm
embodied-reasoning
chain-of-thought
grpo
reinforcement-learning
ontology
benchmark
cosmos
robotics
2026年7月16日
Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
world-model
diffusion
dit
controlnet
multimodal-control
sim2real
autonomous-driving
robotics
physical-ai
video-generation
2026年7月16日
Odyssey-1: A Playable World Model
world-model
interactive-video
real-time-generation
action-conditioned
autoregressive
gaussian-splatting
odyssey
ex-wayve
generative-simulation
2026年7月16日
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
vlm
benchmark
physical-reasoning
embodied-ai
intuitive-physics
robotic-manipulation
iclr2025
multimodal
2026年7月16日
Do generative video models understand physical principles?
world-model
video-generation
benchmark
physical-understanding
evaluation
intuitive-physics
Sora
VideoPoet
2026年7月16日
PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
world-model
video-generation
benchmark
physical-realism
evaluation
anti-physics
text-to-video
MLLM-evaluator
2026年7月16日
PlayerOne: Egocentric World Simulator
world-model
egocentric-video
human-motion-control
video-diffusion-transformer
part-disentangled-motion-injection
4d-reconstruction
ego-exo-dataset
causal-distillation
wan-2-1
2026年7月16日
Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models
jepa
world-model
offline-rl
planning
mpc
latent-dynamics
goal-conditioned
trajectory-stitching
lecun
vicreg
2026年7月16日
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
driving-world-model
policy-world-model
action-free-forecasting
collaborative-state-action-prediction
context-guided-tokenizer
dynamic-focal-loss
autoregressive-transformer
end-to-end-planning
nuscenes
navsim
neurips2025
2026年7月16日
Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
driving-world-model
synthetic-data
3d-asset-insertion
controlnet
dit
nuscenes
corner-case
data-augmentation
fair-evaluation
iclr2026
2026年7月16日
RoboScape: Physics-informed Embodied World Model
world-model
embodied-ai
physics-informed
depth-prediction
keypoint-dynamics
autoregressive-transformer
magvit
agibot-world
synthetic-data
policy-evaluation
2026年7月16日
Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving
world-model
autonomous-driving
3d-occupancy
semi-supervised-learning
volume-rendering
nerf
4d-forecasting
motion-planning
nuscenes
occ3d
2026年7月16日
Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World Model
world-model
interactive-video
autoregressive-diffusion
self-forcing
distillation
kv-cache
real-time
game
action-conditioning
unreal-engine
2026年7月16日
Matrix-Game: Interactive World Foundation Model
world-model
game-generation
minecraft
interactive
image-to-world
diffusion-transformer
action-control
benchmark
open-source
2026年7月16日
SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries
world-model
autonomous-driving
4d-occupancy
sparse-query
occupancy-forecasting
motion-planning
nuscenes
occ3d
aaai2026
2026年7月16日
SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
world-model
autonomous-driving
4d-occupancy
sparse-query
trajectory-conditioning
transformer
feed-forward
nuscenes
occ3d
chamfer-distance
2026年7月16日
The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
world-model
survey
autonomous-driving
taxonomy
driving-world-model
occupancy
point-cloud
video-generation
benchmark
2026年7月16日
Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
world-model
game-generation
interactive-video
camera-control
action-conditioning
autoregressive
diffusion-transformer
distillation
aaa-games
open-source
2026年7月16日
Yan: Foundational Interactive Video Generation
world-model
game-generation
interactive-video
diffusion-forcing
dit
video-diffusion
real-time-simulation
video-editing
action-conditioning
autoregressive-generation
2026年7月16日
TesserAct: Learning 4D Embodied World Models
world-model
4d-scene-reconstruction
rgb-depth-normal
video-diffusion
cogvideox
action-conditioned-prediction
robotic-manipulation
inverse-dynamics
point-cloud
iccv2025
2026年7月16日
Learning Transformer-based World Models with Contrastive Predictive Coding (TWISTER)
world-model
transformer-state-space-model
contrastive-predictive-coding
action-conditioning
atari-100k
deepmind-control-suite
model-based-rl
iclr-2025
2026年7月16日
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
world-model
jepa
self-supervised
video
action-conditioned
zero-shot-planning
robot-manipulation
predictive-embedding
non-generative
2026年7月16日
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
world-model
video-generation
benchmark
evaluation
human-alignment
physics
commonsense
vbench
t2v
2026年7月16日
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
world-model
video-diffusion
robot-manipulation
bimanual-manipulation
inverse-dynamics
cross-embodiment
test-time-scaling
rectified-flow
aloha-robot
2026年7月16日
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
world-model
video-generation
physical-commonsense
benchmark
evaluation
auto-evaluator
action-centric
text-to-video
iclr-2026
2026年7月16日
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
video-wm
latent-dynamics-model
autoregressive-video
magvit-v2
fsq-quantizer
go-benchmark
robotic-manipulation
calvin
rlbench
cvpr-2025
2026年7月16日
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation
world-model
video-generation
physics-aware
text-to-video
mixture-of-experts
cogvideox
wan2.1
physical-commonsense
dataset
lora
2026年7月16日
Marble: A Multimodal World Model
world-model
3d-gaussian-splatting
panorama
text-to-3d
image-to-3d
spatial-intelligence
marble
world-labs
2026年7月16日
RTFM: A Real-Time Frame Model
world-model
real-time-rendering
autoregressive-diffusion-transformer
spatial-memory
persistent-3d
kv-cache
world-labs
rtfm
2026年7月16日
World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
world-model
autonomous-driving
end-to-end-planning
latent-world-model
self-supervised
vision-foundation-model
multi-modal-trajectory
nuscenes
navsim
iccv2025
2026年7月16日
WorldMem: Long-term Consistent World Simulation with Memory
world-model
memory-mechanism
diffusion-forcing
minecraft
long-horizon-consistency
plucker-embedding
cross-attention-memory
neurips-2025
2026年7月16日
WorldModelBench: Judging Video Generation Models As World Models
world-model
video-generation
benchmark
physics-adherence
instruction-following
VLM-judge
human-annotation
reward-gradient
T2V
I2V
2026年7月16日
WorldScore: A Unified Evaluation Benchmark for World Generation
world-model
benchmark
evaluation
video-generation
3d-scene-generation
4d-generation
camera-control
controllability
stanford
iccv-2025
2026年7月16日
WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
world-model
autonomous-driving
gaussian-splatting
feed-forward-reconstruction
latent-diffusion
controlnet
novel-view-synthesis
4d-scene-generation
rectified-flow
iclr2026
2026年7月16日
WoW: Towards a World omniscient World model Through Embodied Interaction
world-model
video-diffusion
embodied-ai
robot-manipulation
diffusion-transformer
inverse-dynamics-model
vlm-critic
physical-reasoning
benchmark
scaling-law
2026年7月16日
Yume: An Interactive World Generation Model
world-model
video-diffusion
image-to-video
camera-control
keyboard-control
masked-diffusion-transformer
diffusion-acceleration
sekai
open-source