AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: reward-model
此标签下有9条笔记。
2026年7月16日
DYNA-1: A Commercial-Grade Autonomous Dexterous Model
vla
reward-model
dexterous-manipulation
autonomous-deployment
napkin-folding
laundry-folding
commercial-robotics
self-recovery
zero-shot-generalization
2026年7月16日
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
benchmark
reward-model
vision-language-model
reinforcement-learning
robot-learning
open-x-embodiment
roboarena
qwen3-vl
data-augmentation
2026年7月16日
Qwen-Image-2.0-RL Technical Report
qwen
rlhf
grpo
flow-grpo
reward-model
on-policy-distillation
reward-hacking
image-editing
cfg
pointwise-reward
post-training
diffusion-rl
2026年7月16日
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
spectrareward
reward-model
t2i-rl
mllm-as-reward
self-reward
unified-multimodal
bagel
prompt-likelihood
training-free
group-relative-rl
awm
geneval
2026年7月16日
Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
driving-world-model
video-diffusion
stable-video-diffusion
action-conditioning
long-horizon-rollout
reward-model
opendv
autonomous-driving
2026年6月25日
Human Preference Score v2 (HPS v2): A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
t2i
human-preference
reward-model
benchmark
clip
rlhf
evaluation
dataset
2026年6月25日
ImageReward & ReFL: Learning and Evaluating Human Preferences for Text-to-Image Generation
t2i
reward-model
rlhf
refl
human-preference
alignment
diffusion
2026年6月25日
SeedEdit 3.0: Fast and High-Quality Generative Image Editing
image-editing
instruction-editing
diffusion
vlm
reward-model
rectified-flow
distillation
bytedance
2026年6月25日
Seedream 3.0 Technical Report
t2i
mmdit
flow-matching
bilingual
text-rendering
high-resolution
reward-model
diffusion-acceleration
repa