AI Research 技术调研

标签: cogvideox

此标签下有6条笔记。

  • 2026年7月16日

    Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction

    • bimanual-manipulation
    • video-prediction
    • optical-flow
    • cogvideox
    • diffusion-policy
    • foundation-policy
    • goal-conditioned-policy
    • text-to-video
  • 2026年7月16日

    Aether: Geometric-Aware Unified World Modeling

    • world-model
    • 4d-reconstruction
    • video-diffusion
    • action-conditioned-prediction
    • visual-planning
    • synthetic-data
    • zero-shot-transfer
    • camera-pose
    • cogvideox
  • 2026年7月16日

    TesserAct: Learning 4D Embodied World Models

    • world-model
    • 4d-scene-reconstruction
    • rgb-depth-normal
    • video-diffusion
    • cogvideox
    • action-conditioned-prediction
    • robotic-manipulation
    • inverse-dynamics
    • point-cloud
    • iccv2025
  • 2026年7月16日

    WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation

    • world-model
    • video-generation
    • physics-aware
    • text-to-video
    • mixture-of-experts
    • cogvideox
    • wan2.1
    • physical-commonsense
    • dataset
    • lora
  • 2026年7月16日

    Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation (WorldDrive)

    • driving-world-model
    • trajectory-vocabulary
    • diffusion-transformer
    • cogvideox
    • representation-inheritance
    • multi-modal-planner
    • future-aware-rewarder
    • navsim
    • nuscenes
  • 2026年6月25日

    视频生成族横向对比(Sora · Veo · Wan · Movie Gen · HunyuanVideo · Kling · CogVideoX 及其谱系)

    • video-generation
    • text-to-video
    • image-to-video
    • diffusion-transformer
    • flow-matching
    • 3d-vae
    • vbench
    • sora
    • veo
    • wan
    • movie-gen
    • hunyuanvideo
    • kling
    • cogvideox
    • deep-dive

  • GitHub