AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: speech
此标签下有7条笔记。
2026年6月25日
AudioLM: a Language Modeling Approach to Audio Generation
audio
speech
music
neural-codec
semantic-tokens
autoregressive
soundstream
w2v-bert
textless-nlp
2026年6月25日
Baichuan-Omni
omni
mllm
audio
video
image
speech
open-source
vita
siglip
whisper
conv-gmlp
2026年6月25日
Qwen2-Audio
audio-language-model
lalm
speech
asr
s2tt
voice-chat
dpo
whisper
qwen
2026年6月25日
Ming-Omni / Ming-Lite-Omni: A Unified Multimodal Model for Perception and Generation
omni
moe
unified
any-to-any
speech
image-generation
video
gpt-4o-class
open-source
2026年6月25日
Qwen2.5-Omni Technical Report
omni
multimodal
any-to-any
speech
thinker-talker
tmrope
streaming
open-source
qwen
2026年6月25日
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
speech
audio-llm
voice-chat
tts
dual-codebook
rlhf
open-source
multimodal
2026年6月25日
Qwen3.5-Omni Technical Report
omni
audio
audio-visual
speech
asr
tts
moe
thinker-talker
streaming
agent
multilingual