AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: reward-free
此标签下有1条笔记。
2026年6月25日
D3PO: Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
diffusion
rlhf
dpo
preference-alignment
reward-free
mdp
text-to-image
cvpr2024