AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: mmdit
此标签下有25条笔记。
2026年6月25日
FLUX1.1 [pro]
t2i
flux
rectified-flow
flow-matching
mmdit
closed-source
api
high-resolution
2026年6月25日
FLUX.1 Tools (Fill / Canny / Depth / Redux)
flux
inpainting
outpainting
controlnet
structural-conditioning
image-variation
mmdit
flow-matching
guidance-distillation
redux
siglip
2026年6月25日
FLUX.1 suite (pro / dev / schnell)
t2i
rectified-flow
flow-matching
mmdit
dit
ladd
guidance-distillation
open-weights
bfl
2026年6月25日
HunyuanVideo: A Systematic Framework For Large Video Generative Models
video
t2v
dit
flow-matching
mmdit
3d-vae
open-source
scaling-law
2026年6月25日
Mochi 1 (preview)
text-to-video
diffusion-transformer
asymmdit
mmdit
3d-attention
flow-matching
open-source
apache-2.0
video-vae
2026年6月25日
SD3.5 Large Turbo 与 (Latent) Adversarial Diffusion Distillation
distillation
adversarial
few-step
mmdit
rectified-flow
turbo
gan
text-to-image
2026年6月25日
Stable Diffusion 3.5 (Large / Large Turbo / Medium)
t2i
mmdit
rectified-flow
open-weights
qk-norm
distillation
sd3
2026年6月25日
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (Stable Diffusion 3 / MMDiT)
t2i
rectified-flow
flow-matching
mmdit
dit
multimodal-transformer
dpo
scaling
open-weights
stability-ai
2026年6月25日
CogView4-6B
t2i
dit
mmdit
flow-matching
glm
bilingual
chinese-text-rendering
open-source
2026年6月25日
FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
flux
in-context-editing
rectified-flow
flow-matching
instruction-editing
character-consistency
mmdit
ladd
open-weights
2026年6月25日
FLUX.2
flux
rectified-flow
mmdit
flow-matching
image-editing
multi-reference
vae
text-rendering
open-weights
2026年6月25日
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
text-to-image
dit
mmdit
moe
sparse-moe
flow-matching
distillation
dmd
gan
open-source
2026年6月25日
HunyuanImage 2.1:高效的 2K 高分辨率文生图扩散模型
t2i
mmdit
diffusion-transformer
high-resolution
2k
distillation
meanflow
rlhf
text-rendering
byt5
vae
repa
open-source
2026年6月25日
Ovis-U1: Unified Understanding, Generation and Editing
unified
mllm
t2i
edit
mmdit
flow-matching
qwen3
ovis
open-source
2026年6月25日
Qwen-Image-Edit / Qwen-Image-Edit-2509
image-editing
mmdit
flow-matching
text-rendering
dual-encoding
qwen2.5-vl
instruction-edit
controlnet
2026年6月25日
Qwen-Image: 阿里巴巴 20B MMDiT 文生图基础模型
t2i
mmdit
text-rendering
image-editing
flow-matching
qwen
chinese-text
open-weights
2026年6月25日
Seedance 1.0: Exploring the Boundaries of Video Generation Models
video
t2v
i2v
dit
mmdit
flow-matching
rlhf
distillation
multi-shot
bytedance
doubao
jimeng
2026年6月25日
Seedream 3.0 Technical Report
t2i
mmdit
flow-matching
bilingual
text-rendering
high-resolution
reward-model
diffusion-acceleration
repa
2026年6月25日
FLUX.2 [klein]
flux
rectified-flow
mmdit
step-distillation
image-editing
multi-reference
apache-2.0
consumer-gpu
Refs
2026年6月25日
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
unified
mllm
mmdit
flow-matching
image-editing
text-rendering
cot
internvl
2026年6月25日
Qwen-Image-2.0 Technical Report
qwen
mmdit
qwen3-vl
vae
rectified-flow
image-editing
text-rendering
rlhf
grpo
dmd-distillation
unified-generation
2026年6月25日
Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling
unified
image-editing
multi-image-composition
hoi
sequence-modeling
mmdit
qwen-image
flow-matching
distillation
dmd
consistency-model
few-step
2026年6月25日
中国系文生图/编辑族横向对比:CogView · Qwen-Image · Hunyuan · Seedream · Kolors · ERNIE 及开源诸侯(2021–2026)
t2i
image-editing
chinese-text-rendering
mmdit
dit
flow-matching
llm-text-encoder
recaption
seedream
qwen-image
cogview
hunyuanimage
ernie
kolors
hidream
lumina
step1x-edit
comparison
2026年6月25日
模型族横向对比:Stable Diffusion → SDXL → SD3 → FLUX 谱系
deep-dive
lineage
stable-diffusion
sdxl
sd3
flux
latent-diffusion
mmdit
rectified-flow
t2i
open-weights
2026年6月25日
模型架构演进:从 U-Net 扩散到统一 omni 骨干(2020–2026)
omni
architecture
diffusion
dit
mmdit
rectified-flow
autoregressive
visual-tokenizer
vae
text-encoder
unified-multimodal
survey