AI Research 技术调研

标签: image-to-text

此标签下有3条笔记。

  • 2026年6月25日

    Versatile Diffusion: Text, Images and Variations All in One Diffusion Model

    • unified
    • multimodal
    • diffusion
    • multi-flow
    • t2i
    • image-to-text
    • image-variation
    • ldm
    • clip
  • 2026年6月25日

    CM3Leon: Scaling Autoregressive Multi-Modal Models (Pretraining and Instruction Tuning)

    • autoregressive
    • token-based
    • retrieval-augmented
    • text-to-image
    • image-to-text
    • instruction-tuning
    • contrastive-decoding
    • cm3
    • chameleon-lineage
  • 2026年6月25日

    SEED: Planting a SEED of Vision in Large Language Model

    • unified
    • visual-tokenizer
    • discrete-tokens
    • vq
    • multimodal-llm
    • autoregressive
    • q-former
    • image-to-text
    • text-to-image

  • GitHub