AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: preference-alignment
此标签下有2条笔记。
2026年6月25日
D3PO: Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
diffusion
rlhf
dpo
preference-alignment
reward-free
mdp
text-to-image
cvpr2024
2026年6月25日
训练方法:目标函数·多阶段·偏好对齐·蒸馏
training
objective
flow-matching
rectified-flow
preference-alignment
rlhf
dpo
grpo
distillation
few-step
consistency
sampler
editing
survey
omni