AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: multimodal
此标签下有21条笔记。
2026年7月16日
Ego4D: Around the World in 3,000 Hours of Egocentric Video
egocentric-video
dataset
benchmark
human-video
embodied-ai
first-person
multimodal
pretraining-corpus
2026年7月16日
Implicit Behavioral Cloning (IBC)
imitation-learning
behavioral-cloning
energy-based-model
implicit-policy
multimodal
manipulation
d4rl
corl2021
2026年7月16日
Learning to Model the World With Language (Dynalang)
world-model
language-grounding
rssm
dreamerv3
multimodal
model-based-rl
imitation-free
text-pretraining
vision-language-navigation
2026年7月16日
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
vlm
benchmark
physical-reasoning
embodied-ai
intuitive-physics
robotic-manipulation
iclr2025
multimodal
2026年6月25日
GauGAN2 / PoE-GAN — 文字 + 语义涂鸦 + 草图多模态合成风景图
gan
multimodal
text-to-image
semantic-image-synthesis
sketch
product-of-experts
nvidia-canvas
spade
2026年6月25日
M6: A Chinese Multimodal Pretrainer
multimodal
chinese
pretraining
moe
text-to-image
vqgan
encoder-decoder
mixture-of-experts
2026年6月25日
CM3: A Causal Masked Multimodal Model of the Internet
autoregressive
decoder-only
causal-masking
multimodal
html
vqvae-gan
zero-shot
infilling
entity-linking
unified
2026年6月25日
UniDiffuser: One Transformer Fits All Distributions in Multi-Modal Diffusion
unified
multimodal
diffusion
transformer
u-vit
t2i
i2t
joint-generation
latent-diffusion
2026年6月25日
Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
unified
multimodal
diffusion
multi-flow
t2i
image-to-text
image-variation
ldm
clip
2026年6月25日
DreamLLM: Synergistic Multimodal Comprehension and Creation
unified
mllm
interleaved
diffusion
score-distillation
image-generation
multimodal
2026年6月25日
VideoPoet: A Large Language Model for Zero-Shot Video Generation
video
autoregressive
llm
discrete-token
multimodal
text-to-video
image-to-video
audio
magvit-v2
zero-shot
2026年6月25日
Aurora (Grok Image Generation)
autoregressive
mixture-of-experts
next-token
interleaved
multimodal
image-generation
image-editing
photorealism
closed-source
grok
2026年6月25日
GPT-4o 原生图像生成 (4o image generation / gpt-image-1)
omni
autoregressive
native-image
image-generation
text-rendering
image-editing
multimodal
closed-source
2026年6月25日
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
unified
multimodal
rectified-flow
autoregression
llm
t2i
vlm
deepseek
2026年6月25日
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
unified
multimodal
diffusion
next-token
transformer
late-fusion
chameleon
vae
2026年6月25日
MMaDA: Multimodal Large Diffusion Language Models
unified
discrete-diffusion
masked-diffusion
dllm
multimodal
t2i
reasoning
rl
grpo
neurips2025
2026年6月25日
Qwen2.5-Omni Technical Report
omni
multimodal
any-to-any
speech
thinker-talker
tmrope
streaming
open-source
qwen
2026年6月25日
Seedream 4.0: Toward Next-generation Multimodal Image Generation
t2i
image-editing
multimodal
dit
high-compression-vae
joint-post-training
rlhf
adversarial-distillation
quantization
speculative-decoding
4k
multi-image-reference
in-context-reasoning
closed-source
2026年6月25日
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
speech
audio-llm
voice-chat
tts
dual-codebook
rlhf
open-source
multimodal
2026年6月25日
可灵 3.0 系列模型(Kling 3.0:Video 3.0 / Video 3.0 Omni / Image 3.0 / Image 3.0 Omni)
video-generation
omni
multimodal
native-audio
multi-shot
character-reference
closed-source
kuaishou
kling
2026年6月25日
Seedance 2.0
video
audio-video
multimodal
t2v
i2v
r2v
editing
binaural-audio
closed-source