AI Research 技术调研

标签: speech

此标签下有7条笔记。

  • 2026年6月25日

    AudioLM: a Language Modeling Approach to Audio Generation

    • audio
    • speech
    • music
    • neural-codec
    • semantic-tokens
    • autoregressive
    • soundstream
    • w2v-bert
    • textless-nlp
  • 2026年6月25日

    Baichuan-Omni

    • omni
    • mllm
    • audio
    • video
    • image
    • speech
    • open-source
    • vita
    • siglip
    • whisper
    • conv-gmlp
  • 2026年6月25日

    Qwen2-Audio

    • audio-language-model
    • lalm
    • speech
    • asr
    • s2tt
    • voice-chat
    • dpo
    • whisper
    • qwen
  • 2026年6月25日

    Ming-Omni / Ming-Lite-Omni: A Unified Multimodal Model for Perception and Generation

    • omni
    • moe
    • unified
    • any-to-any
    • speech
    • image-generation
    • video
    • gpt-4o-class
    • open-source
  • 2026年6月25日

    Qwen2.5-Omni Technical Report

    • omni
    • multimodal
    • any-to-any
    • speech
    • thinker-talker
    • tmrope
    • streaming
    • open-source
    • qwen
  • 2026年6月25日

    Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

    • speech
    • audio-llm
    • voice-chat
    • tts
    • dual-codebook
    • rlhf
    • open-source
    • multimodal
  • 2026年6月25日

    Qwen3.5-Omni Technical Report

    • omni
    • audio
    • audio-visual
    • speech
    • asr
    • tts
    • moe
    • thinker-talker
    • streaming
    • agent
    • multilingual

  • GitHub