AI Research 技术调研

标签: magvit-v2

此标签下有4条笔记。

  • 2026年7月16日

    VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

    • video-wm
    • latent-dynamics-model
    • autoregressive-video
    • magvit-v2
    • fsq-quantizer
    • go-benchmark
    • robotic-manipulation
    • calvin
    • rlbench
    • cvpr-2025
  • 2026年6月25日

    VideoPoet: A Large Language Model for Zero-Shot Video Generation

    • video
    • autoregressive
    • llm
    • discrete-token
    • multimodal
    • text-to-video
    • image-to-video
    • audio
    • magvit-v2
    • zero-shot
  • 2026年6月25日

    W.A.L.T: Photorealistic Video Generation with Diffusion Models

    • video-diffusion
    • latent-diffusion
    • transformer
    • dit
    • window-attention
    • causal-3d-vae
    • text-to-video
    • magvit-v2
  • 2026年6月25日

    Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

    • unified
    • understanding-generation
    • discrete-diffusion
    • maskgit
    • autoregressive
    • magvit-v2
    • phi-1.5

  • GitHub