AI Research 技术调研

标签: training

此标签下有2条笔记。

  • 2026年7月16日

    世界模型训练方法(想象力 RL · 扩散/流 · next-token · JEPA 回归 · 蒸馏)

    • world-model
    • training
    • imagination-rl
    • dreamer
    • muzero
    • diffusion
    • flow-matching
    • diffusion-forcing
    • autoregressive
    • jepa
    • distillation
  • 2026年6月25日

    训练方法:目标函数·多阶段·偏好对齐·蒸馏

    • training
    • objective
    • flow-matching
    • rectified-flow
    • preference-alignment
    • rlhf
    • dpo
    • grpo
    • distillation
    • few-step
    • consistency
    • sampler
    • editing
    • survey
    • omni

  • GitHub