AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: evaluation
此标签下有12条笔记。
2026年7月16日
具身智能评测(LIBERO/SimplerEnv/CALVIN/Meta-World · 真机成功率 · 泛化)
embodied
benchmark
evaluation
manipulation
generalization
sim-to-real
real2sim
libero
calvin
simpler-env
metaworld
maniskill
reproducibility
2026年7月16日
VBench: Comprehensive Benchmark Suite for Video Generative Models
world-model
video-generation
benchmark
evaluation
human-alignment
t2v
vbench
cvpr2024
2026年7月16日
VideoPhy: Evaluating Physical Commonsense for Video Generation
world-model
video-generation
physical-commonsense
benchmark
evaluation
auto-evaluator
videocon-physics
text-to-video
iclr-2025
2026年7月16日
Do generative video models understand physical principles?
world-model
video-generation
benchmark
physical-understanding
evaluation
intuitive-physics
Sora
VideoPoet
2026年7月16日
PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
world-model
video-generation
benchmark
physical-realism
evaluation
anti-physics
text-to-video
MLLM-evaluator
2026年7月16日
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
world-model
video-generation
benchmark
evaluation
human-alignment
physics
commonsense
vbench
t2v
2026年7月16日
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
world-model
video-generation
physical-commonsense
benchmark
evaluation
auto-evaluator
action-centric
text-to-video
iclr-2026
2026年7月16日
WorldScore: A Unified Evaluation Benchmark for World Generation
world-model
benchmark
evaluation
video-generation
3d-scene-generation
4d-generation
camera-control
controllability
stanford
iccv-2025
2026年7月16日
世界模型评测(物理一致性 · 3D 一致性 · 动作保真 · 下游 RL/规划 · 驾驶指标)
world-model
benchmark
evaluation
physics
controllability
planning
autonomous-driving
超人
2026年6月25日
Human Preference Score v2 (HPS v2): A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
t2i
human-preference
reward-model
benchmark
clip
rlhf
evaluation
dataset
2026年6月25日
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
t2i
benchmark
evaluation
judge-model
mllm-judge
creator-centric
qwen
2026年6月25日
评测 benchmark 演进与横向数字
benchmark
evaluation
fid
clipscore
geneval
dpg-bench
t2i-compbench
hpsv2
imagereward
pickscore
mjhq-30k
arena-elo
lmarena
ocr-text-render
gedit
magicbrush
imgedit
vbench
movie-gen-bench
wise
omni
survey