AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: audio
此标签下有10条笔记。
2026年7月16日
A-JEPA: Joint-Embedding Predictive Architecture Can Listen
jepa
audio
self-supervised
representation-learning
spectrogram
masking
curriculum-learning
vision-transformer
2026年7月16日
Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
jepa
world-model
self-supervised
music
audio
representation-learning
stem-separation
non-generative
retrieval
ismir
2026年6月25日
AudioLM: a Language Modeling Approach to Audio Generation
audio
speech
music
neural-codec
semantic-tokens
autoregressive
soundstream
w2v-bert
textless-nlp
2026年6月25日
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model
any-modality
multimodal-llm
llama-2
perceiver-resampler
frozen-llm
audio
video
imu
instruction-tuning
understanding
2026年6月25日
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
audio
text-to-audio
text-to-music
text-to-speech
latent-diffusion
self-supervised
audiomae
gpt-2
unified
2026年6月25日
VideoPoet: A Large Language Model for Zero-Shot Video Generation
video
autoregressive
llm
discrete-token
multimodal
text-to-video
image-to-video
audio
magvit-v2
zero-shot
2026年6月25日
Baichuan-Omni
omni
mllm
audio
video
image
speech
open-source
vita
siglip
whisper
conv-gmlp
2026年6月25日
Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
audio
music-generation
latent-diffusion
text-to-audio
timing-conditioning
stereo
vae
clap
2026年6月25日
Sora 2
video
audio
text-to-video
world-model
diffusion
synced-audio
cameo
closed-source
2026年6月25日
Qwen3.5-Omni Technical Report
omni
audio
audio-visual
speech
asr
tts
moe
thinker-talker
streaming
agent
multilingual