AI Research 技术调研

标签: audio-visual

此标签下有3条笔记。

  • 2026年7月16日

    ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data

    • audio-visual
    • contact-microphone
    • ear-in-hand
    • diffusion-policy
    • umi
    • in-the-wild-data
    • imitation-learning
    • contact-rich-manipulation
    • corl-2024
  • 2026年7月16日

    Wan-Streamer v0.2: Higher Resolution, Same Latency

    • wan-streamer
    • real-time
    • full-duplex
    • audio-visual
    • streaming-video-generation
    • digital-human
    • block-causal-attention
    • flow-matching
    • ulysses
    • context-parallel
    • kv-cache
    • low-latency
  • 2026年6月25日

    Qwen3.5-Omni Technical Report

    • omni
    • audio
    • audio-visual
    • speech
    • asr
    • tts
    • moe
    • thinker-talker
    • streaming
    • agent
    • multilingual

  • GitHub