AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: t2i
此标签下有77条笔记。
2026年7月16日
Seedream 5.0 Pro
t2i
image-editing
unified-multimodal
interactive-editing
grounding
bbox
layer-separation
infographic
text-rendering
photographic-realism
multilingual
closed-source
2026年6月25日
DALL·E: Zero-Shot Text-to-Image Generation
autoregressive
transformer
dvae
vqvae
t2i
zero-shot
sparse-attention
clip-rerank
2026年6月25日
DALL·E 2 / unCLIP(Hierarchical Text-Conditional Image Generation with CLIP Latents)
t2i
diffusion
unclip
clip
prior-decoder
guidance
openai
2026年6月25日
Imagen: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
t2i
diffusion
cascaded-diffusion
t5
classifier-free-guidance
drawbench
pixel-space
2026年6月25日
Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors
t2i
autoregressive
vqgan
scene-control
segmentation
classifier-free-guidance
human-prior
meta
2026年6月25日
Midjourney (公测 V1–V4)
t2i
closed-source
diffusion
discord
aesthetic
midjourney
niji
commercial
2026年6月25日
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
t2i
diffusion
mixture-of-experts
space-moe
time-moe
latent-diffusion
edge-supervised
coco-fid
2026年6月25日
All are Worth Words: A ViT Backbone for Diffusion Models (U-ViT)
diffusion
vit
backbone
transformer
long-skip-connection
latent-diffusion
t2i
class-conditional
Layers
Heads
Params
2026年6月25日
UniDiffuser: One Transformer Fits All Distributions in Multi-Modal Diffusion
unified
multimodal
diffusion
transformer
u-vit
t2i
i2t
joint-generation
latent-diffusion
2026年6月25日
Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
unified
multimodal
diffusion
multi-flow
t2i
image-to-text
image-variation
ldm
clip
2026年6月25日
Adobe Firefly(首个图像模型 + Photoshop Generative Fill)
t2i
diffusion
commercial-safe
generative-fill
editing
closed-source
adobe-stock
content-credentials
2026年6月25日
CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images
t2i
latent-diffusion
creative-commons
synthetic-caption
copyright
blip-2
stable-diffusion
data-efficient
2026年6月25日
DALL·E 3
t2i
diffusion
latent-diffusion
recaptioning
synthetic-caption
prompt-following
chatgpt
t5
2026年6月25日
Human Preference Score v2 (HPS v2): A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
t2i
human-preference
reward-model
benchmark
clip
rlhf
evaluation
dataset
2026年6月25日
InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation
t2i
rectified-flow
reflow
distillation
one-step
flow-matching
acceleration
stable-diffusion
2026年6月25日
Kandinsky 2.x (Image Prior + Latent Diffusion)
t2i
unclip
image-prior
latent-diffusion
movq
multilingual
open-source
clip
2026年6月25日
Midjourney V5 / V5.1 / V5.2
t2i
diffusion
closed-source
commercial
photorealism
outpainting
style-tuner
discord
2026年6月25日
Muse: Text-To-Image Generation via Masked Generative Transformers
t2i
masked-generative
transformer
vqgan
parallel-decoding
maskgit
t5
discrete-tokens
2026年6月25日
PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
t2i
dit
diffusion-transformer
cross-attention
t5
low-cost-training
re-captioning
open-source
Params
Images
2026年6月25日
Playground v2 (1024px Aesthetic)
t2i
diffusion
sdxl
latent-diffusion
aesthetic
open-weights
mjhq-30k
2026年6月25日
ImageReward & ReFL: Learning and Evaluating Human Preferences for Text-to-Image Generation
t2i
reward-model
rlhf
refl
human-preference
alignment
diffusion
2026年6月25日
SDXL-Turbo / Adversarial Diffusion Distillation (ADD)
t2i
distillation
diffusion
gan
few-step
real-time
sdxl
score-distillation
2026年6月25日
StyleDrop: Text-to-Image Generation in Any Style
style-tuning
personalization
few-shot
adapter
peft
muse
masked-transformer
t2i
feedback
2026年6月25日
Würstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion Models
t2i
latent-diffusion
cascade
compression
vqgan
efficient-training
convnext
stable-cascade
2026年6月25日
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
t2i
relay-diffusion
cascaded-diffusion
latent-diffusion
distillation
recaption
unet
dit
cogview
2026年6月25日
DALL·E 3 系列产品化(ChatGPT 集成 / API)
t2i
latent-diffusion
recaptioning
prompt-following
chatgpt
closed-source
openai
safety
system-card
2026年6月25日
ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
t2i
diffusion
llm-text-encoder
adapter
prompt-following
dpg-bench
resampler
adaln
sd15
sdxl
training-free-unet
可训练参数
2026年6月25日
Emu3: Next-Token Prediction is All You Need
unified
autoregressive
next-token-prediction
discrete-token
vision-tokenizer
t2i
t2v
vlm
movqgan
dpo
2026年6月25日
FLUX1.1 [pro]
t2i
flux
rectified-flow
flow-matching
mmdit
closed-source
api
high-resolution
2026年6月25日
FLUX.1 suite (pro / dev / schnell)
t2i
rectified-flow
flow-matching
mmdit
dit
ladd
guidance-distillation
open-weights
bfl
2026年6月25日
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
autoregressive
visual-tokenizer
residual-diffusion
var
t2i
efficient-inference
1024px
Params
Step
2026年6月25日
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
t2i
diffusion-transformer
dit
chinese
bilingual
multi-resolution
rope
recaptioning
mllm
open-source
2026年6月25日
Ideogram 2.0
t2i
typography
text-rendering
closed-source
diffusion
product-launch
2026年6月25日
Imagen 3
t2i
latent-diffusion
google
gemini
closed-source
synthetic-caption
human-eval
synthid
2026年6月25日
Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
t2i
autoregressive
var
next-scale-prediction
bitwise-tokenizer
bsq
infinite-vocabulary
self-correction
scaling-law
2026年6月25日
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
unified
multimodal
rectified-flow
autoregression
llm
t2i
vlm
deepseek
2026年6月25日
Kolors(可图): Effective Training of Diffusion Model for Photorealistic Text-to-Image Synthesis
t2i
latent-diffusion
unet
sdxl
chatglm
bilingual
chinese-text-rendering
recaption
open-source
2026年6月25日
Liquid: Language Models are Scalable and Unified Multi-modal Generators
unified-multimodal
autoregressive
vqgan
discrete-token
next-token-prediction
scaling-law
text-to-image
t2i
mllm
2026年6月25日
LlamaGen: Autoregressive Model Beats Diffusion — Llama for Scalable Image Generation
autoregressive
next-token
image-tokenizer
vqgan
llama
t2i
class-conditional
vllm
2026年6月25日
Midjourney V6 / V6.1
t2i
diffusion
aesthetics
closed-source
text-rendering
midjourney
2026年6月25日
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
t2i
dit
diffusion-transformer
4k
kv-compression
weak-to-strong
efficient
pixart
Params
2026年6月25日
Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models
t2i
diffusion
dit
deep-fusion
llm-text-encoder
llama3
graphic-design
text-rendering
capsbench
closed-source
2026年6月25日
Recraft V3 (red_panda)
t2i
closed-source
text-rendering
layout-control
vector-graphics
controlnet
design
api
ocr
2026年6月25日
Sana: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
t2i
diffusion-transformer
linear-attention
deep-compression-ae
rectified-flow
efficient
on-device
gemma
2026年6月25日
Stable Diffusion 3.5 (Large / Large Turbo / Medium)
t2i
mmdit
rectified-flow
open-weights
qk-norm
distillation
sd3
2026年6月25日
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (Stable Diffusion 3 / MMDiT)
t2i
rectified-flow
flow-matching
mmdit
dit
multimodal-transformer
dpo
scaling
open-weights
stability-ai
2026年6月25日
CogView4-6B
t2i
dit
mmdit
flow-matching
glm
bilingual
chinese-text-rendering
open-source
2026年6月25日
Adobe Firefly Image Model 4 / 4 Ultra
t2i
firefly
adobe
commercially-safe
closed-source
creative-cloud
photorealism
2026年6月25日
Imagen 4 Fast / Imagen 4(Ultra) GA 与 GPT Image 1 mini
t2i
closed-source
api
cost-tier
imagen-4
gpt-image-1
autoregressive
latent-diffusion
synthid
c2pa
2026年6月25日
HunyuanImage 2.1:高效的 2K 高分辨率文生图扩散模型
t2i
mmdit
diffusion-transformer
high-resolution
2k
distillation
meanflow
rlhf
text-rendering
byt5
vae
repa
open-source
2026年6月25日
HunyuanImage 3.0 Technical Report
t2i
unified-multimodal
autoregressive
moe
diffusion
transfusion
chain-of-thought
open-source
rlhf
2026年6月25日
Ideogram 3.0
t2i
typography
text-rendering
style-reference
graphic-design
closed-source
diffusion
product-launch
inpaint
fine-tuning
2026年6月25日
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
unified-multimodal
autoregressive
decoupled-visual-encoding
t2i
vlm
open-source
deepseek
2026年6月25日
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
discrete-diffusion
unified-multimodal
masked-generation
dllm
t2i
image-editing
image-understanding
self-grpo
open-source
2026年6月25日
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
t2i
dit
flow-matching
unified-attention
gemma2
recaptioning
open-source
iccv2025
2026年6月25日
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
autoregressive
t2i
unified-generation
decoder-only
vqgan
image-editing
controllable-generation
speculative-jacobi
2026年6月25日
Midjourney V7
t2i
closed-source
personalization
draft-mode
omni-reference
midjourney
aesthetics
2026年6月25日
MMaDA: Multimodal Large Diffusion Language Models
unified
discrete-diffusion
masked-diffusion
dllm
multimodal
t2i
reasoning
rl
grpo
neurips2025
2026年6月25日
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
autoregressive
continuous-tokens
flow-matching
next-token-prediction
t2i
image-editing
qwen2.5
flux-vae
nextstep-grpo
flowgrpo
2026年6月25日
OmniGen2: Towards Instruction-Aligned Multimodal Generation
unified
multimodal-generation
t2i
image-editing
in-context
decoupled-decoding
omni-rope
grpo
rectified-flow
open-source
2026年6月25日
Ovis-U1: Unified Understanding, Generation and Editing
unified
mllm
t2i
edit
mmdit
flow-matching
qwen3
ovis
open-source
2026年6月25日
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
t2i
linear-attention
dit
model-growth
depth-pruning
inference-scaling
came-8bit
efficient-scaling
geneval
sana
2026年6月25日
Qwen-Image: 阿里巴巴 20B MMDiT 文生图基础模型
t2i
mmdit
text-rendering
image-editing
flow-matching
qwen
chinese-text
open-weights
2026年6月25日
Reve Image 1.0 (Halfmoon)
t2i
closed-source
prompt-adherence
typography
intent-driven
palo-alto
startup
halfmoon
2026年6月25日
Seedream 3.0 Technical Report
t2i
mmdit
flow-matching
bilingual
text-rendering
high-resolution
reward-model
diffusion-acceleration
repa
2026年6月25日
Seedream 4.0: Toward Next-generation Multimodal Image Generation
t2i
image-editing
multimodal
dit
high-compression-vae
joint-post-training
rlhf
adversarial-distillation
quantization
speculative-decoding
4k
multi-image-reference
in-context-reasoning
closed-source
2026年6月25日
Z-Image / Z-Image-Turbo:单流扩散 Transformer 的高效 6B 文生图基础模型
t2i
diffusion-transformer
single-stream
dmd
distillation
rlhf
bilingual-text
image-editing
efficient
2026年6月25日
Ideogram 4.0
t2i
open-weight
dit
single-stream
flow-matching
qwen3-vl
json-prompt
text-rendering
bounding-box
2026年6月25日
Krea 2
t2i
diffusion-transformer
flow-matching
dit
style-reference
moodboard
open-weights
distillation
rl
dpo
krea
2026年6月25日
Midjourney V8.1
t2i
midjourney
closed-source
aesthetic-preference
hd
fast-inference
2026年6月25日
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
t2i
benchmark
evaluation
judge-model
mllm-judge
creator-centric
qwen
2026年6月25日
Qwen-Image-Flash: Beyond Objective Design
few-step-distillation
dmd
flow-matching
t2i
image-editing
distillation-recipe
qwen-image
2026年6月25日
Recraft V4
t2i
closed-source
design
vector-graphics
svg
raster
text-rendering
aesthetics
api
no-tech-report
2026年6月25日
Reve 2.0
t2i
layout
image-editing
4k
diffusion
llm-planner
qwen
agentic
closed-source
2026年6月25日
Seedream 5.0 Lite
t2i
image-editing
unified-multimodal
deep-thinking
visual-reasoning
online-search
web-search
in-context-reasoning
world-knowledge
information-visualization
multi-image-reference
4k
closed-source
elo
magicbench
2026年6月25日
中国系文生图/编辑族横向对比:CogView · Qwen-Image · Hunyuan · Seedream · Kolors · ERNIE 及开源诸侯(2021–2026)
t2i
image-editing
chinese-text-rendering
mmdit
dit
flow-matching
llm-text-encoder
recaption
seedream
qwen-image
cogview
hunyuanimage
ernie
kolors
hidream
lumina
step1x-edit
comparison
2026年6月25日
模型族横向对比:Stable Diffusion → SDXL → SD3 → FLUX 谱系
deep-dive
lineage
stable-diffusion
sdxl
sd3
flux
latent-diffusion
mmdit
rectified-flow
t2i
open-weights