AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: mllm
此标签下有13条笔记。
2026年7月16日
RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
vla
mllm
embodied-brain
task-planning
affordance-perception
trajectory-prediction
sharerobot
lora
open-x-embodiment
cvpr2025
2026年7月16日
ADriver-I: A General World Model for Autonomous Driving
world-model
autonomous-driving
mllm
video-diffusion
action-conditioning
vicuna
nuscenes
closed-loop
control-signal-prediction
2026年6月25日
DreamLLM: Synergistic Multimodal Comprehension and Creation
unified
mllm
interleaved
diffusion
score-distillation
image-generation
multimodal
2026年6月25日
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
mllm
subject-driven
personalization
zero-shot
image-as-foreign-language
score-distillation
alignernet
kosmos
stable-diffusion
2026年6月25日
MGIE: Guiding Instruction-based Image Editing via Multimodal Large Language Models
instruction-editing
mllm
llava
diffusion
instructpix2pix
expressive-instruction
apple
2026年6月25日
Baichuan-Omni
omni
mllm
audio
video
image
speech
open-source
vita
siglip
whisper
conv-gmlp
2026年6月25日
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
t2i
diffusion-transformer
dit
chinese
bilingual
multi-resolution
rope
recaptioning
mllm
open-source
2026年6月25日
Liquid: Language Models are Scalable and Unified Multi-modal Generators
unified-multimodal
autoregressive
vqgan
discrete-token
next-token-prediction
scaling-law
text-to-image
t2i
mllm
2026年6月25日
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
unified
mllm
instruction-tuning
vpit
continuous-visual-tokens
autoregressive
diffusion-autoencoder
llama-3
2026年6月25日
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
unified
mllm
external-diffuser
sdxl
image-editing
llama2
vit-bridge
comprehension-generation
2026年6月25日
Ovis-U1: Unified Understanding, Generation and Editing
unified
mllm
t2i
edit
mmdit
flow-matching
qwen3
ovis
open-source
2026年6月25日
Step1X-Edit: A Practical Framework for General Image Editing
image-editing
instruction-editing
mllm
dit
flux
qwen2.5-vl
rectified-flow
gedit-bench
open-source
2026年6月25日
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
unified
mllm
mmdit
flow-matching
image-editing
text-rendering
cot
internvl