AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: unified
此标签下有32条笔记。
2026年6月25日
ERNIE-ViLG: Unified Generative Pre-training for Bidirectional Vision-Language Generation
text-to-image
image-captioning
autoregressive
vqgan
unified
chinese
baidu
2026年6月25日
CM3: A Causal Masked Multimodal Model of the Internet
autoregressive
decoder-only
causal-masking
multimodal
html
vqvae-gan
zero-shot
infilling
entity-linking
unified
2026年6月25日
UniDiffuser: One Transformer Fits All Distributions in Multi-Modal Diffusion
unified
multimodal
diffusion
transformer
u-vit
t2i
i2t
joint-generation
latent-diffusion
2026年6月25日
Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
unified
multimodal
diffusion
multi-flow
t2i
image-to-text
image-variation
ldm
clip
2026年6月25日
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
audio
text-to-audio
text-to-music
text-to-speech
latent-diffusion
self-supervised
audiomae
gpt-2
unified
2026年6月25日
DreamLLM: Synergistic Multimodal Comprehension and Creation
unified
mllm
interleaved
diffusion
score-distillation
image-generation
multimodal
2026年6月25日
Emu: Generative Pretraining in Multimodality
unified
autoregressive
lmm
multimodal-generation
interleaved
video-text
eva-clip
llama
stable-diffusion
in-context-learning
2026年6月25日
GILL: Generating Images with Multimodal Language Models
unified
frozen-llm
interleaved
retrieval
image-generation
mapping-network
opt
stable-diffusion
neurips-2023
2026年6月25日
SEED: Planting a SEED of Vision in Large Language Model
unified
visual-tokenizer
discrete-tokens
vq
multimodal-llm
autoregressive
q-former
image-to-text
text-to-image
2026年6月25日
Chameleon: Mixed-Modal Early-Fusion Foundation Models
unified
early-fusion
token-based
autoregressive
mixed-modal
vqgan
image-tokenizer
multimodal-llm
2026年6月25日
Emu3: Next-Token Prediction is All You Need
unified
autoregressive
next-token-prediction
discrete-token
vision-tokenizer
t2i
t2v
vlm
movqgan
dpo
2026年6月25日
Gemini 2.0 Flash 原生图像生成(Native Image Output)
native-image-output
interleaved-text-image
unified
multimodal-llm
conversational-editing
synthid
closed-source
nano-banana-lineage
2026年6月25日
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
unified
autoregressive
vqgan
siglip
decoupled-encoder
any-to-any
deepseek
2026年6月25日
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
unified
multimodal
rectified-flow
autoregression
llm
t2i
vlm
deepseek
2026年6月25日
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
unified
mllm
instruction-tuning
vpit
continuous-visual-tokens
autoregressive
diffusion-autoencoder
llama-3
2026年6月25日
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
unified
mllm
external-diffuser
sdxl
image-editing
llama2
vit-bridge
comprehension-generation
2026年6月25日
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
unified
understanding-generation
discrete-diffusion
maskgit
autoregressive
magvit-v2
phi-1.5
2026年6月25日
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
unified
multimodal
diffusion
next-token
transformer
late-fusion
chameleon
vae
2026年6月25日
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
unified
autoregressive
next-token
visual-tokenizer
rq-vae
clip-alignment
vlm
image-generation
video-generation
2026年6月25日
DreamO: A Unified Framework for Image Customization
image-customization
dit
flux
lora
subject-driven
identity
try-on
style
feature-routing
unified
2026年6月25日
Gemini 2.0 Flash 原生图像生成(公开实验版 gemini-2.0-flash-exp,2025-03)
gemini
native-image-output
interleaved-text-image
conversational-editing
unified
multimodal-llm
world-knowledge
text-rendering
synthid
closed-source
nano-banana-lineage
2026年6月25日
Ming-Omni / Ming-Lite-Omni: A Unified Multimodal Model for Perception and Generation
omni
moe
unified
any-to-any
speech
image-generation
video
gpt-4o-class
open-source
2026年6月25日
MMaDA: Multimodal Large Diffusion Language Models
unified
discrete-diffusion
masked-diffusion
dllm
multimodal
t2i
reasoning
rl
grpo
neurips2025
2026年6月25日
OmniGen2: Towards Instruction-Aligned Multimodal Generation
unified
multimodal-generation
t2i
image-editing
in-context
decoupled-decoding
omni-rope
grpo
rectified-flow
open-source
2026年6月25日
Ovis-U1: Unified Understanding, Generation and Editing
unified
mllm
t2i
edit
mmdit
flow-matching
qwen3
ovis
open-source
2026年6月25日
Show-o2: Improved Native Unified Multimodal Models
unified
understanding-generation
autoregressive
flow-matching
3d-causal-vae
video
qwen2.5
omni-attention
2026年6月25日
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
unified
edit
perception
siglip
flux
qwen2.5-vl
flow-matching
open-source
2026年6月25日
VACE: All-in-One Video Creation and Editing
video
editing
unified
dit
controllable
inpainting
reference-to-video
wan
ltx-video
iccv2025
2026年6月25日
X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
unified
autoregressive
discrete-token
rl
grpo
text-rendering
image-generation
image-understanding
tokenizer
siglip
flux
2026年6月25日
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
unified
mllm
mmdit
flow-matching
image-editing
text-rendering
cot
internvl
2026年6月25日
Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling
unified
image-editing
multi-image-composition
hoi
sequence-modeling
mmdit
qwen-image
flow-matching
distillation
dmd
consistency-model
few-step
2026年6月25日
统一理解生成 & any-to-any 全模态专题
unified
understanding-generation
any-to-any
omni
early-fusion
decoupled-encoder
diffusion-ar-hybrid
autoregressive
rectified-flow
thinker-talker