AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: unified-multimodal
此标签下有12条笔记。
2026年7月16日
Seedream 5.0 Pro
t2i
image-editing
unified-multimodal
interactive-editing
grounding
bbox
layer-separation
infographic
text-rendering
photographic-realism
multilingual
closed-source
2026年7月16日
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
spectrareward
reward-model
t2i-rl
mllm-as-reward
self-reward
unified-multimodal
bagel
prompt-likelihood
training-free
group-relative-rl
awm
geneval
2026年7月16日
Doe-1: Closed-Loop Autonomous Driving with Large World Model
driving-world-model
closed-loop-autonomous-driving
autoregressive-transformer
chameleon
lumina-mgpt
action-tokenization
unified-multimodal
nuscenes
next-token-prediction
2026年6月25日
Liquid: Language Models are Scalable and Unified Multi-modal Generators
unified-multimodal
autoregressive
vqgan
discrete-token
next-token-prediction
scaling-law
text-to-image
t2i
mllm
2026年6月25日
BAGEL: Emerging Properties in Unified Multimodal Pretraining
unified-multimodal
mot
mixture-of-transformers
rectified-flow
interleaved
world-modeling
image-editing
open-source
2026年6月25日
BLIP3-o: A Family of Fully Open Unified Multimodal Models
unified-multimodal
clip-feature-diffusion
flow-matching
diffusion-transformer
sequential-training
fully-open
image-generation
image-understanding
2026年6月25日
HunyuanImage 3.0 Technical Report
t2i
unified-multimodal
autoregressive
moe
diffusion
transfusion
chain-of-thought
open-source
rlhf
2026年6月25日
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
unified-multimodal
autoregressive
decoupled-visual-encoding
t2i
vlm
open-source
deepseek
2026年6月25日
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
discrete-diffusion
unified-multimodal
masked-generation
dllm
t2i
image-editing
image-understanding
self-grpo
open-source
2026年6月25日
Seedream 5.0 Lite
t2i
image-editing
unified-multimodal
deep-thinking
visual-reasoning
online-search
web-search
in-context-reasoning
world-knowledge
information-visualization
multi-image-reference
4k
closed-source
elo
magicbench
2026年6月25日
统一/Omni 模型族横向对比:Chameleon · Emu · Janus · BAGEL · OmniGen · Show-o · VAR 谱系
unified-multimodal
omni
autoregressive
diffusion
flow-matching
next-scale-prediction
understanding-generation
image-editing
deep-dive
2026年6月25日
模型架构演进:从 U-Net 扩散到统一 omni 骨干(2020–2026)
omni
architecture
diffusion
dit
mmdit
rectified-flow
autoregressive
visual-tokenizer
vae
text-encoder
unified-multimodal
survey