AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: dit
此标签下有62条笔记。
2026年7月16日
OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
world-model
autonomous-driving
occupancy-generation
diffusion-transformer
4d-occupancy
trajectory-conditioning
nuscenes
vq-vae
dit
2026年7月16日
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
world-model
autonomous-driving
synthetic-data
controlnet
multi-view
lidar-generation
diffusion
dit
cosmos
data-flywheel
2026年7月16日
History-Guided Video Diffusion
world-model
video-diffusion
diffusion-forcing
history-guidance
classifier-free-guidance
dit
long-video-generation
autoregressive-rollout
icml-2025
2026年7月16日
Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
world-model
diffusion
dit
controlnet
multimodal-control
sim2real
autonomous-driving
robotics
physical-ai
video-generation
2026年7月16日
Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
driving-world-model
synthetic-data
3d-asset-insertion
controlnet
dit
nuscenes
corner-case
data-augmentation
fair-evaluation
iclr2026
2026年7月16日
Yan: Foundational Interactive Video Generation
world-model
game-generation
interactive-video
diffusion-forcing
dit
video-diffusion
real-time-simulation
video-editing
action-conditioning
autoregressive-generation
2026年6月25日
Scalable Diffusion Models with Transformers (DiT)
dit
diffusion-transformer
latent-diffusion
adaln-zero
scaling
imagenet
backbone
vit
2026年6月25日
GenTron: Diffusion Transformers for Image and Video Generation
dit
diffusion-transformer
text-to-image
text-to-video
scaling
motion-free-guidance
t2i-compbench
Param
2026年6月25日
PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
t2i
dit
diffusion-transformer
cross-attention
t5
low-cost-training
re-captioning
open-source
Params
Images
2026年6月25日
W.A.L.T: Photorealistic Video Generation with Diffusion Models
video-diffusion
latent-diffusion
transformer
dit
window-attention
causal-3d-vae
text-to-video
magvit-v2
2026年6月25日
Allegro: Open the Black Box of Commercial-Level Video Generation Model
text-to-video
dit
video-vae
3d-rope
full-attention
open-weights
flow-free-diffusion
rhymes-ai
2026年6月25日
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
text-to-video
diffusion-transformer
3d-vae
expert-adaln
dit
open-source
i2v
2026年6月25日
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
t2i
relay-diffusion
cascaded-diffusion
latent-diffusion
distillation
recaption
unet
dit
cogview
2026年6月25日
FLUX.1 suite (pro / dev / schnell)
t2i
rectified-flow
flow-matching
mmdit
dit
ladd
guidance-distillation
open-weights
bfl
2026年6月25日
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
t2i
diffusion-transformer
dit
chinese
bilingual
multi-resolution
rope
recaptioning
mllm
open-source
2026年6月25日
HunyuanVideo: A Systematic Framework For Large Video Generative Models
video
t2v
dit
flow-matching
mmdit
3d-vae
open-source
scaling-law
2026年6月25日
In-Context LoRA for Diffusion Transformers (IC-LoRA)
in-context
lora
dit
flux
image-set
task-agnostic
editing
customization
2026年6月25日
Kling (可灵) 视频生成大模型
video-generation
t2v
i2v
dit
3d-vae
spatiotemporal-attention
closed-source
kuaishou
2026年6月25日
LTX-Video: Realtime Video Latent Diffusion
video
t2v
i2v
dit
latent-diffusion
rectified-flow
video-vae
realtime
open-source
2026年6月25日
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
dit
flow-matching
flag-dit
text-to-image
text-to-video
text-to-3d
text-to-speech
rope
resolution-extrapolation
unified-generation
2026年6月25日
OminiControl: Minimal and Universal Control for Diffusion Transformer
dit
flux
controlnet
subject-driven
image-conditioning
lora
rope
dataset
2026年6月25日
Open-Sora Plan / Open-Sora(2024 开源复现 Sora)
video
t2v
i2v
dit
diffusion
sora-reproduction
open-source
wf-vae
skiparse-attention
rectified-flow
2026年6月25日
Sora(Sora Turbo)公开发布
video-generation
text-to-video
diffusion-transformer
dit
spacetime-patches
world-simulator
recaptioning
closed-source
2026年6月25日
PixArt-δ: Fast and Controllable Image Generation with Latent Consistency Models
pixart
dit
lcm
consistency-distillation
controlnet
transformer
text-to-image
few-step
huawei
2026年6月25日
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
t2i
dit
diffusion-transformer
4k
kv-compression
weak-to-strong
efficient
pixart
Params
2026年6月25日
Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models
t2i
diffusion
dit
deep-fusion
llm-text-encoder
llama3
graphic-design
text-rendering
capsbench
closed-source
2026年6月25日
Pyramid Flow:金字塔式流匹配的高效视频生成
video-generation
flow-matching
pyramidal-flow
autoregressive
dit
mm-dit
efficient-training
2026年6月25日
REPA: Representation Alignment for Generation — Training Diffusion Transformers Is Easier Than You Think
diffusion-transformer
representation-alignment
dinov2
sit
dit
self-supervised
training-efficiency
flow-matching
imagenet
2026年6月25日
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
tts
speech-synthesis
zero-shot
voice-cloning
autoregressive
diffusion
dit
rl
voice-conversion
in-context-learning
2026年6月25日
SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
diffusion
flow-matching
stochastic-interpolant
dit
imagenet
t2i-backbone
sde
ode
2026年6月25日
Sora: Video generation models as world simulators
video-generation
text-to-video
diffusion-transformer
dit
spacetime-patches
world-simulator
recaptioning
native-resolution
scaling
closed-source
2026年6月25日
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (Stable Diffusion 3 / MMDiT)
t2i
rectified-flow
flow-matching
mmdit
dit
multimodal-transformer
dpo
scaling
open-weights
stability-ai
2026年6月25日
TRELLIS: Structured 3D Latents for Scalable and Versatile 3D Generation
3d-generation
slat
rectified-flow
dit
sparse-voxel
dinov2
image-to-3d
text-to-3d
gaussian-splatting
mesh
2026年6月25日
ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling
instruction-edit
subject-reference
flux
inpainting
lora
diffusion
rectified-flow
dit
2026年6月25日
CogView4-6B
t2i
dit
mmdit
flow-matching
glm
bilingual
chinese-text-rendering
open-source
2026年6月25日
Cosmos World Foundation Model Platform for Physical AI (Cosmos-Predict1)
world-model
physical-ai
video-generation
diffusion
autoregressive
dit
tokenizer
robotics
autonomous-driving
open-weight
2026年6月25日
DreamO: A Unified Framework for Image Customization
image-customization
dit
flux
lora
subject-driven
identity
try-on
style
feature-routing
unified
2026年6月25日
FLUX.1 Krea [dev]
text-to-image
rectified-flow
dit
flux
post-training
rlhf
dpo
aesthetics
photorealism
guidance-distillation
open-weights
2026年6月25日
Goku: Flow Based Video Generative Foundation Models
video-generation
text-to-image
image-to-video
rectified-flow
dit
joint-image-video
3d-vae
bytedance
2026年6月25日
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
text-to-image
dit
mmdit
moe
sparse-moe
flow-matching
distillation
dmd
gan
open-source
2026年6月25日
Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material
image-to-3d
shape-generation
pbr-texture
flow-matching
dit
vae
multi-view-diffusion
open-source
tencent
2026年6月25日
HunyuanVideo 1.5
video-generation
dit
t2v
i2v
flow-matching
sparse-attention
sparse-attn
video-super-resolution
open-source
lightweight
muon
distillation
2026年6月25日
ICEdit:用 In-Context 生成范式做指令图像编辑(Enabling Instructional Image Editing with In-Context Generation in Large-Scale DiT)
image-editing
instruction-editing
in-context
flux-fill
lora
moe
inference-time-scaling
dit
rectified-flow
neurips2025
2026年6月25日
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
video-generation
text-to-video
image-to-video
text-to-image
image-editing
latent-diffusion
flow-matching
dit
crossdit
nabla
sparse-attention
open-source
russian
2026年6月25日
可灵 Kling 2.0 / 2.1 / 2.5 Turbo
video-generation
t2v
i2v
dit
closed-source
kuaishou
kling
mvl
video-editing
2026年6月25日
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
t2i
dit
flow-matching
unified-attention
gemma2
recaptioning
open-source
iccv2025
2026年6月25日
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
t2i
linear-attention
dit
model-growth
depth-pruning
inference-scaling
came-8bit
efficient-scaling
geneval
sana
2026年6月25日
Diffusion Transformers with Representation Autoencoders (RAE)
diffusion-transformer
representation-encoder
dinov2
autoencoder
latent-diffusion
imagenet
flow-matching
dit
2026年6月25日
Seedance 1.0: Exploring the Boundaries of Video Generation Models
video
t2v
i2v
dit
mmdit
flow-matching
rlhf
distillation
multi-shot
bytedance
doubao
jimeng
2026年6月25日
Seedream 4.0: Toward Next-generation Multimodal Image Generation
t2i
image-editing
multimodal
dit
high-compression-vae
joint-post-training
rlhf
adversarial-distillation
quantization
speculative-decoding
4k
multi-image-reference
in-context-reasoning
closed-source
2026年6月25日
SkyReels-V2: Infinite-Length Film Generative Model
video
diffusion-forcing
autoregressive
flow-matching
dit
dpo
long-video
open-source
2026年6月25日
Step1X-Edit: A Practical Framework for General Image Editing
image-editing
instruction-editing
mllm
dit
flux
qwen2.5-vl
rectified-flow
gedit-bench
open-source
2026年6月25日
UNO: Less-to-More Generalization — Unlocking More Controllability by In-Context Generation
subject-driven
customization
in-context-generation
dit
flux
lora
rope
multi-subject
data-synthesis
2026年6月25日
VACE: All-in-One Video Creation and Editing
video
editing
unified
dit
controllable
inpainting
reference-to-video
wan
ltx-video
iccv2025
2026年6月25日
Wan: Open and Advanced Large-Scale Video Generative Models (Wan 2.1)
video
t2v
i2v
dit
flow-matching
3d-vae
wan-vae
open-source
video-editing
vace
2026年6月25日
ERNIE-Image Technical Report
text-to-image
dit
single-stream
latent-diffusion
flow-matching
dpo
distillation
dmd
text-rendering
open-weights
chinese
2026年6月25日
Ideogram 4.0
t2i
open-weight
dit
single-stream
flow-matching
qwen3-vl
json-prompt
text-rendering
bounding-box
2026年6月25日
Krea 2
t2i
diffusion-transformer
flow-matching
dit
style-reference
moodboard
open-weights
distillation
rl
dpo
krea
2026年6月25日
Qwen-Image-VAE-2.0 Technical Report
vae
high-compression
latent-diffusion
tokenizer
text-rendering
semantic-alignment
dit
image-generation
2026年6月25日
Seed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content Generation
3d-generation
image-to-3d
pbr
vecset
rectified-flow
dit
moe
vae
articulation
scene-generation
simulation-ready
bytedance
2026年6月25日
中国系文生图/编辑族横向对比:CogView · Qwen-Image · Hunyuan · Seedream · Kolors · ERNIE 及开源诸侯(2021–2026)
t2i
image-editing
chinese-text-rendering
mmdit
dit
flow-matching
llm-text-encoder
recaption
seedream
qwen-image
cogview
hunyuanimage
ernie
kolors
hidream
lumina
step1x-edit
comparison
2026年6月25日
模型架构演进:从 U-Net 扩散到统一 omni 骨干(2020–2026)
omni
architecture
diffusion
dit
mmdit
rectified-flow
autoregressive
visual-tokenizer
vae
text-encoder
unified-multimodal
survey