AI Research 技术调研

标签: text-to-speech

此标签下有2条笔记。

  • 2026年6月25日

    AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

    • audio
    • text-to-audio
    • text-to-music
    • text-to-speech
    • latent-diffusion
    • self-supervised
    • audiomae
    • gpt-2
    • unified
  • 2026年6月25日

    Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

    • dit
    • flow-matching
    • flag-dit
    • text-to-image
    • text-to-video
    • text-to-3d
    • text-to-speech
    • rope
    • resolution-extrapolation
    • unified-generation

  • GitHub