AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: rlhf
此标签下有15条笔记。
2026年7月16日
Qwen-Image-2.0-RL Technical Report
qwen
rlhf
grpo
flow-grpo
reward-model
on-policy-distillation
reward-hacking
image-editing
cfg
pointwise-reward
post-training
diffusion-rl
2026年6月25日
D3PO: Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
diffusion
rlhf
dpo
preference-alignment
reward-free
mdp
text-to-image
cvpr2024
2026年6月25日
DDPO: Training Diffusion Models with Reinforcement Learning
diffusion
rl
rlhf
rlaif
policy-gradient
ppo
mdp
text-to-image
reward-optimization
vlm-feedback
2026年6月25日
Diffusion-DPO: Diffusion Model Alignment Using Direct Preference Optimization
diffusion
dpo
alignment
rlhf
preference-optimization
sdxl
text-to-image
2026年6月25日
Human Preference Score v2 (HPS v2): A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
t2i
human-preference
reward-model
benchmark
clip
rlhf
evaluation
dataset
2026年6月25日
ImageReward & ReFL: Learning and Evaluating Human Preferences for Text-to-Image Generation
t2i
reward-model
rlhf
refl
human-preference
alignment
diffusion
2026年6月25日
FLUX.1 Krea [dev]
text-to-image
rectified-flow
dit
flux
post-training
rlhf
dpo
aesthetics
photorealism
guidance-distillation
open-weights
2026年6月25日
HunyuanImage 2.1:高效的 2K 高分辨率文生图扩散模型
t2i
mmdit
diffusion-transformer
high-resolution
2k
distillation
meanflow
rlhf
text-rendering
byt5
vae
repa
open-source
2026年6月25日
HunyuanImage 3.0 Technical Report
t2i
unified-multimodal
autoregressive
moe
diffusion
transfusion
chain-of-thought
open-source
rlhf
2026年6月25日
Seedance 1.0: Exploring the Boundaries of Video Generation Models
video
t2v
i2v
dit
mmdit
flow-matching
rlhf
distillation
multi-shot
bytedance
doubao
jimeng
2026年6月25日
Seedream 4.0: Toward Next-generation Multimodal Image Generation
t2i
image-editing
multimodal
dit
high-compression-vae
joint-post-training
rlhf
adversarial-distillation
quantization
speculative-decoding
4k
multi-image-reference
in-context-reasoning
closed-source
2026年6月25日
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
speech
audio-llm
voice-chat
tts
dual-codebook
rlhf
open-source
multimodal
2026年6月25日
Z-Image / Z-Image-Turbo:单流扩散 Transformer 的高效 6B 文生图基础模型
t2i
diffusion-transformer
single-stream
dmd
distillation
rlhf
bilingual-text
image-editing
efficient
2026年6月25日
Qwen-Image-2.0 Technical Report
qwen
mmdit
qwen3-vl
vae
rectified-flow
image-editing
text-rendering
rlhf
grpo
dmd-distillation
unified-generation
2026年6月25日
训练方法:目标函数·多阶段·偏好对齐·蒸馏
training
objective
flow-matching
rectified-flow
preference-alignment
rlhf
dpo
grpo
distillation
few-step
consistency
sampler
editing
survey
omni