AI Research 技术调研

标签: any-to-any

此标签下有7条笔记。

  • 2026年7月16日

    Nano Banana 2 Lite(Gemini 3.1 Flash Lite Image)与 Gemini Omni Flash 开发者 GA

    • nano-banana
    • gemini-3-1-flash-lite-image
    • gemini-omni-flash
    • any-to-any
    • video-generation
    • conversational-editing
    • interactions-api
    • synthid
    • closed-source
    • tpu
  • 2026年6月25日

    Omni / 多模态生成技术演进调研 (2020 → 2026-07) · 主汇总

    • omni
    • multimodal-generation
    • text-to-image
    • image-editing
    • unified-understanding-generation
    • any-to-any
    • video-generation
    • diffusion
    • autoregressive
    • survey
  • 2026年6月25日

    NExT-GPT: Any-to-Any Multimodal LLM

    • any-to-any
    • mm-llm
    • omni
    • diffusion-decoder
    • imagebind
    • vicuna
    • instruction-tuning
    • mosit
    • projection-layer
  • 2026年6月25日

    Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

    • unified
    • autoregressive
    • vqgan
    • siglip
    • decoupled-encoder
    • any-to-any
    • deepseek
  • 2026年6月25日

    Ming-Omni / Ming-Lite-Omni: A Unified Multimodal Model for Perception and Generation

    • omni
    • moe
    • unified
    • any-to-any
    • speech
    • image-generation
    • video
    • gpt-4o-class
    • open-source
  • 2026年6月25日

    Qwen2.5-Omni Technical Report

    • omni
    • multimodal
    • any-to-any
    • speech
    • thinker-talker
    • tmrope
    • streaming
    • open-source
    • qwen
  • 2026年6月25日

    统一理解生成 & any-to-any 全模态专题

    • unified
    • understanding-generation
    • any-to-any
    • omni
    • early-fusion
    • decoupled-encoder
    • diffusion-ar-hybrid
    • autoregressive
    • rectified-flow
    • thinker-talker

  • GitHub