AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: benchmark
此标签下有62条笔记。
2026年7月16日
Habitat: A Platform for Embodied AI Research
embodied-ai
simulator
pointgoal-navigation
photorealistic
matterport3d
gibson
reinforcement-learning
benchmark
sim2real
2026年7月16日
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
benchmark
meta-rl
multi-task-rl
robotic-manipulation
mujoco
sawyer
reinforcement-learning
generalization
2026年7月16日
RLBench: The Robot Learning Benchmark & Learning Environment
robot-learning
manipulation
benchmark
simulation
few-shot
imitation-learning
coppeliasim
pyrep
franka-panda
ra-l
2026年7月16日
SAPIEN: A SimulAted Part-based Interactive ENvironment
simulator
articulated-objects
partnet-mobility
physx
ros
manipulation
sim-to-real
embodied-ai
part-level
benchmark
类别
模型
可动部件
2026年7月16日
CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
benchmark
language-conditioned
manipulation
long-horizon
imitation-learning
play-data
pybullet
franka
2026年7月16日
Ego4D: Around the World in 3,000 Hours of Egocentric Video
egocentric-video
dataset
benchmark
human-video
embodied-ai
first-person
multimodal
pretraining-corpus
2026年7月16日
ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
sim-infra
manipulation
benchmark
sapien
partnet-mobility
point-cloud
learning-from-demonstrations
reinforcement-learning
generalization
2026年7月16日
Orbit: A Unified Simulation Framework for Interactive Robot Learning Environments
orbit
isaac-lab
isaac-sim
physx-5
gpu-simulation
sim-to-real
robot-learning
deformable-body
benchmark
rl-infra
2026年7月16日
LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
libero
benchmark
lifelong-learning
robot-manipulation
imitation-learning
behavioral-cloning
procedural-generation
pddl
declarative-vs-procedural
vla-eval
2026年7月16日
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
maniskill2
benchmark
sim-infra
sapien
mpm-soft-body
render-server
gpu-simulation
imitation-learning
reinforcement-learning
sim2real
2026年7月16日
RoboVQA: Multimodal Long-Horizon Reasoning for Robotics
robovqa
benchmark
vqa
long-horizon-planning
cross-embodiment
intervention-rate
video-language-model
videococa
chain-of-thought
data-collection
2026年7月16日
BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
embodied-ai
benchmark
simulation
mobile-manipulation
omnigibson
long-horizon
household
sim-to-real
bddl
2026年7月16日
BiGym: A Demo-Driven Mobile Bi-Manual Manipulation Benchmark
benchmark
embodied-ai
mobile-manipulation
bimanual-manipulation
humanoid
mujoco
unitree-h1
imitation-learning
demo-driven-rl
vr-teleoperation
2026年7月16日
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
embodied-ai
benchmark
rlbench
generalization
vision-language-manipulation
point-cloud-transformer
llm-planning
vlm-grounding
3d-policy
imitation-learning
2026年7月16日
GRUtopia: Dream General Robots in a City at Scale
embodied-ai
simulation
isaac-sim
llm-npc
city-scale
scene-dataset
benchmark
navigation
mobile-manipulation
2026年7月16日
HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation
embodied-ai
benchmark
humanoid
whole-body-control
loco-manipulation
mujoco
unitree-h1
dexterous-hands
hierarchical-rl
tactile
2026年7月16日
ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
sim-infra
gpu-simulation
parallel-rendering
robot-learning
manipulation
sim2real
real2sim
sapien
benchmark
rl
2026年7月16日
MetaUrban: An Embodied AI Simulation Platform for Urban Micromobility
sim-infra
urban-navigation
micromobility
procedural-generation
embodied-ai
digital-humans
safe-rl
imitation-learning
benchmark
2026年7月16日
Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER)
benchmark
real2sim
manipulation
vla
policy-evaluation
sapien
maniskill
google-robot
widowx
bridge
2026年7月16日
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
benchmark
robotic-manipulation
generalization
distribution-shift
rlbench
coppeliasim
pyrep
sim-to-real
behavior-cloning
rss-2024
2026年7月16日
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
benchmark
vla
long-horizon-reasoning
mujoco
simulation
vlm-evaluation
language-conditioned-manipulation
common-sense
dataset
generalization
2026年7月16日
AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
embodied-ai
manipulation-dataset
bimanual
humanoid
VLA
latent-action
teleoperation
GO-1
benchmark
2026年7月16日
EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
embodied-ai
dexterous-manipulation
egocentric-video
hand-tracking
Apple-Vision-Pro
ARKit
imitation-learning
manipulation-dataset
benchmark
human-video
2026年7月16日
A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
survey
world-models
physical-simulators
embodied-intelligence
humanoid-robot
autonomous-driving
sim2real
robot-taxonomy
benchmark
model-based-rl
2026年7月16日
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
benchmark
mllm-agent
embodied-ai
ai2-thor
habitat
vlmbench
alfred
vision-language-action
capability-oriented-evaluation
icml-2025
2026年7月16日
GenManip: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
benchmark
isaac-sim
scene-graph
llm-task-generation
tabletop-manipulation
modular-vlm-agent
vla
behavior-cloning
cvpr2025
2026年7月16日
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
embodied-ai
benchmark
generalist-policy
real-robot-eval
pairwise-comparison
bradley-terry
elo
droid
chatbot-arena
crowd-sourced
2026年7月16日
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
benchmark
long-horizon-manipulation
vla
system1-system2
hierarchical-planning
libero
openvla
vlm-planner
neurips-2025
2026年7月16日
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
benchmark
dual-arm-manipulation
bimanual-manipulation
domain-randomization
sim-to-real
mllm-code-generation
cross-embodiment
vla-training-data
2026年7月16日
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
benchmark
dual-arm-manipulation
bimanual-manipulation
digital-twin
sim2real
llm-code-generation
3d-generative-model
maniskill
imitation-learning
CVPR2025
2026年7月16日
SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending
humanoid
whole-body-control
loco-manipulation
hierarchical-rl
skill-blending
goal-conditioned
cross-embodiment
benchmark
reward-hacking
isaac-gym
2026年7月16日
Survey of Vision-Language-Action Models for Embodied Manipulation
survey
vision-language-action
embodied-manipulation
robot-learning
imitation-learning
reinforcement-learning
action-tokenization
chain-of-thought
hierarchical-vla
benchmark
2026年7月16日
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
benchmark
reward-model
vision-language-model
reinforcement-learning
robot-learning
open-x-embodiment
roboarena
qwen3-vl
data-augmentation
2026年7月16日
具身智能评测(LIBERO/SimplerEnv/CALVIN/Meta-World · 真机成功率 · 泛化)
embodied
benchmark
evaluation
manipulation
generalization
sim-to-real
real2sim
libero
calvin
simpler-env
metaworld
maniskill
reproducibility
2026年7月16日
Physion: Evaluating Physical Prediction from Vision in Humans and Machines
world-model
intuitive-physics
benchmark
physical-prediction
human-comparison
object-centric
graph-neural-network
threedworld
neurips2021
2026年7月16日
Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications
world-model
autonomous-driving
occupancy-forecasting
benchmark
camera-only
4d-occupancy
nuscenes
lyft-level5
cvpr2024
2026年7月16日
LightZero: A Unified Benchmark for Monte Carlo Tree Search in General Sequential Decision Scenarios
world-model
mcts
muzero
alphazero
model-based-rl
benchmark
open-source-toolkit
tree-search
reinforcement-learning
2026年7月16日
Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties
world-model
intuitive-physics
benchmark
physical-property-inference
latent-properties
threedworld
neurips2023
human-comparison
object-centric
dpi-net
2026年7月16日
VBench: Comprehensive Benchmark Suite for Video Generative Models
world-model
video-generation
benchmark
evaluation
human-alignment
t2v
vbench
cvpr2024
2026年7月16日
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
world-model
video-generation
physical-commonsense
benchmark
vlm-evaluator
text-to-video
hierarchical-evaluation
phygenbench
phygeneval
2026年7月16日
VideoPhy: Evaluating Physical Commonsense for Video Generation
world-model
video-generation
physical-commonsense
benchmark
evaluation
auto-evaluator
videocon-physics
text-to-video
iclr-2025
2026年7月16日
WorldSimBench: Towards Video Generation Models as World Simulators
world-model
video-generation
benchmark
world-simulator
human-preference-evaluator
embodied-evaluation
autonomous-driving
robot-manipulation
minecraft
closed-loop-evaluation
指令
视频
维度
动作类别
正例
负例
2026年7月16日
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
world-model
robotic-manipulation
video-diffusion
vla
flow-matching
dual-arm
cross-embodiment
neural-simulator
benchmark
2026年7月16日
Impossible Videos
world-model
video-generation
video-understanding
benchmark
counterfactual
physical-commonsense
Video-LLM
taxonomy
ICML2025
2026年7月16日
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
physical-ai
vlm
embodied-reasoning
chain-of-thought
grpo
reinforcement-learning
ontology
benchmark
cosmos
robotics
2026年7月16日
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
vlm
benchmark
physical-reasoning
embodied-ai
intuitive-physics
robotic-manipulation
iclr2025
multimodal
2026年7月16日
Do generative video models understand physical principles?
world-model
video-generation
benchmark
physical-understanding
evaluation
intuitive-physics
Sora
VideoPoet
2026年7月16日
PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
world-model
video-generation
benchmark
physical-realism
evaluation
anti-physics
text-to-video
MLLM-evaluator
2026年7月16日
Matrix-Game: Interactive World Foundation Model
world-model
game-generation
minecraft
interactive
image-to-world
diffusion-transformer
action-control
benchmark
open-source
2026年7月16日
The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
world-model
survey
autonomous-driving
taxonomy
driving-world-model
occupancy
point-cloud
video-generation
benchmark
2026年7月16日
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
world-model
video-generation
benchmark
evaluation
human-alignment
physics
commonsense
vbench
t2v
2026年7月16日
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
world-model
video-generation
physical-commonsense
benchmark
evaluation
auto-evaluator
action-centric
text-to-video
iclr-2026
2026年7月16日
WorldModelBench: Judging Video Generation Models As World Models
world-model
video-generation
benchmark
physics-adherence
instruction-following
VLM-judge
human-annotation
reward-gradient
T2V
I2V
2026年7月16日
WorldScore: A Unified Evaluation Benchmark for World Generation
world-model
benchmark
evaluation
video-generation
3d-scene-generation
4d-generation
camera-control
controllability
stanford
iccv-2025
2026年7月16日
WoW: Towards a World omniscient World model Through Embodied Interaction
world-model
video-diffusion
embodied-ai
robot-manipulation
diffusion-transformer
inverse-dynamics-model
vlm-critic
physical-reasoning
benchmark
scaling-law
2026年7月16日
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
world-model
robot-policy-evaluation
benchmark
video-diffusion
autoregressive-generation
action-conditioning
dmd2-distillation
vlm-judge
wan
2026年7月16日
世界模型评测(物理一致性 · 3D 一致性 · 动作保真 · 下游 RL/规划 · 驾驶指标)
world-model
benchmark
evaluation
physics
controllability
planning
autonomous-driving
超人
2026年6月25日
Emu Edit: Precise Image Editing via Recognition and Generation Tasks
instruction-editing
image-editing
multi-task
task-embedding
diffusion
benchmark
emu
2026年6月25日
Human Preference Score v2 (HPS v2): A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
t2i
human-preference
reward-model
benchmark
clip
rlhf
evaluation
dataset
2026年6月25日
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
image-editing
instruction-guided
dataset
benchmark
multi-turn
dall-e-2
instructpix2pix
neurips-2023
2026年6月25日
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
t2i
benchmark
evaluation
judge-model
mllm-judge
creator-centric
qwen
2026年6月25日
评测 benchmark 演进与横向数字
benchmark
evaluation
fid
clipscore
geneval
dpg-bench
t2i-compbench
hpsv2
imagereward
pickscore
mjhq-30k
arena-elo
lmarena
ocr-text-render
gedit
magicbrush
imgedit
vbench
movie-gen-bench
wise
omni
survey