AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: text-to-image
此标签下有46条笔记。
2026年7月16日
Scaling Robot Learning with Semantically Imagined Experience
data-augmentation
diffusion-inpainting
imagen-editor
owl-vit
rt-1
text-to-image
distractor-robustness
success-detection
sim-free-scaling
llm-prompting
2026年7月16日
Cosmos-Predict2: A Suite of Diffusion-based World Foundation Models Available in 2B, and 14B
world-model
physical-ai
video-generation
text-to-image
diffusion-transformer
robotics
autonomous-driving
natten
action-conditioned
cosmos
2026年7月16日
JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
jepa
text-to-image
cross-attention
flow-matching
masked-prediction
imagenet
multimodal-fusion
clip
mar
non-generative-lineage
2026年6月25日
Omni / 多模态生成技术演进调研 (2020 → 2026-07) · 主汇总
omni
multimodal-generation
text-to-image
image-editing
unified-understanding-generation
any-to-any
video-generation
diffusion
autoregressive
survey
2026年6月25日
CLIP-Guided Diffusion (Katherine Crowson / 社区版)
clip-guidance
classifier-guidance
diffusion
text-to-image
training-free
community
ablated-diffusion
2026年6月25日
CogView: Mastering Text-to-Image Generation via Transformers
autoregressive
vqvae
transformer
gpt
text-to-image
chinese
fp16-stability
sandwich-ln
pb-relax
2026年6月25日
ERNIE-ViLG: Unified Generative Pre-training for Bidirectional Vision-Language Generation
text-to-image
image-captioning
autoregressive
vqgan
unified
chinese
baidu
2026年6月25日
GauGAN2 / PoE-GAN — 文字 + 语义涂鸦 + 草图多模态合成风景图
gan
multimodal
text-to-image
semantic-image-synthesis
sketch
product-of-experts
nvidia-canvas
spade
2026年6月25日
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
diffusion
classifier-free-guidance
clip-guidance
text-to-image
inpainting
adm
openai
2026年6月25日
High-Resolution Image Synthesis with Latent Diffusion Models (LDM)
latent-diffusion
ldm
vae
u-net
cross-attention
text-to-image
stable-diffusion
two-stage
2026年6月25日
M6: A Chinese Multimodal Pretrainer
multimodal
chinese
pretraining
moe
text-to-image
vqgan
encoder-decoder
mixture-of-experts
2026年6月25日
NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion
unified-generation
autoregressive
vq-gan
3d-transformer
sparse-attention
text-to-image
text-to-video
image-editing
any-to-vision
2026年6月25日
VQ-Diffusion:用于文生图的向量量化扩散模型
discrete-diffusion
mask-and-replace
vq-vae
text-to-image
non-autoregressive
masked-generation
2026年6月25日
VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance
vqgan
clip
text-to-image
clip-guidance
training-free
image-editing
ai-art
latent-optimization
eleutherai
2026年6月25日
eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers
diffusion
text-to-image
ensemble-of-experts
cascaded-diffusion
t5
clip
paint-with-words
edm
nvidia
2026年6月25日
ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-Experts
text-to-image
diffusion
latent-diffusion
mixture-of-experts
knowledge-enhanced
chinese
baidu
cvpr2023
params
图
2026年6月25日
Parti: Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
autoregressive
text-to-image
vit-vqgan
seq2seq
scaling
classifier-free-guidance
tpu
partiprompts
2026年6月25日
Prompt-to-Prompt Image Editing with Cross Attention Control
diffusion
image-editing
cross-attention
training-free
text-to-image
attention-injection
imagen
stable-diffusion
2026年6月25日
Stable Diffusion v1 (CompVis / Stability AI / Runway)
latent-diffusion
text-to-image
open-weights
ldm
unet
clip
laion
2026年6月25日
Stable Diffusion 2.0 / 2.1
latent-diffusion
text-to-image
open-weights
ldm
unet
openclip
v-prediction
upscaler
depth2img
laion
2026年6月25日
CM3Leon: Scaling Autoregressive Multi-Modal Models (Pretraining and Instruction Tuning)
autoregressive
token-based
retrieval-augmented
text-to-image
image-to-text
instruction-tuning
contrastive-decoding
cm3
chameleon-lineage
2026年6月25日
D3PO: Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
diffusion
rlhf
dpo
preference-alignment
reward-free
mdp
text-to-image
cvpr2024
2026年6月25日
DDPO: Training Diffusion Models with Reinforcement Learning
diffusion
rl
rlhf
rlaif
policy-gradient
ppo
mdp
text-to-image
reward-optimization
vlm-feedback
2026年6月25日
DeepFloyd IF
text-to-image
cascaded-diffusion
pixel-diffusion
t5-xxl
imagen-style
open-source
text-rendering
super-resolution
2026年6月25日
Diffusion-DPO: Diffusion Model Alignment Using Direct Preference Optimization
diffusion
dpo
alignment
rlhf
preference-optimization
sdxl
text-to-image
2026年6月25日
One-step Diffusion with Distribution Matching Distillation (DMD)
diffusion-distillation
one-step
distribution-matching
vsd
score-distillation
text-to-image
acceleration
2026年6月25日
Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack
text-to-image
latent-diffusion
quality-tuning
aesthetic-alignment
sft
meta
2026年6月25日
GenTron: Diffusion Transformers for Image and Video Generation
dit
diffusion-transformer
text-to-image
text-to-video
scaling
motion-free-guidance
t2i-compbench
Param
2026年6月25日
Ideogram 0.1
text-to-image
typography
text-rendering
closed-source
commercial
imagen-team
diffusion
2026年6月25日
Imagen 2
text-to-image
diffusion
closed-source
google-deepmind
vertex-ai
synthid
inpainting
outpainting
aesthetics-conditioning
recaptioning
2026年6月25日
Latent Consistency Models (LCM / LCM-LoRA)
diffusion
distillation
consistency-model
few-step
lora
acceleration
text-to-image
stable-diffusion
2026年6月25日
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
latent-diffusion
text-to-image
unet
dual-text-encoder
refiner
micro-conditioning
multi-aspect
open-weights
2026年6月25日
SEED: Planting a SEED of Vision in Large Language Model
unified
visual-tokenizer
discrete-tokens
vq
multimodal-llm
autoregressive
q-former
image-to-text
text-to-image
2026年6月25日
StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis
gan
stylegan
text-to-image
fast-inference
clip
dino
single-step
2026年6月25日
Liquid: Language Models are Scalable and Unified Multi-modal Generators
unified-multimodal
autoregressive
vqgan
discrete-token
next-token-prediction
scaling-law
text-to-image
t2i
mllm
2026年6月25日
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
text-to-image
diffusion-transformer
next-dit
flow-matching
rectified-flow
rope
resolution-extrapolation
few-step-sampling
multilingual
unified-generation
2026年6月25日
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
dit
flow-matching
flag-dit
text-to-image
text-to-video
text-to-3d
text-to-speech
rope
resolution-extrapolation
unified-generation
2026年6月25日
PixArt-δ: Fast and Controllable Image Generation with Latent Consistency Models
pixart
dit
lcm
consistency-distillation
controlnet
transformer
text-to-image
few-step
huawei
2026年6月25日
SD3.5 Large Turbo 与 (Latent) Adversarial Diffusion Distillation
distillation
adversarial
few-step
mmdit
rectified-flow
turbo
gan
text-to-image
2026年6月25日
Stable Cascade (Würstchen v3)
text-to-image
latent-diffusion
cascade
wuerstchen
efficient
convnext
vqgan
semantic-compressor
open-weights
2026年6月25日
FLUX.1 Krea [dev]
text-to-image
rectified-flow
dit
flux
post-training
rlhf
dpo
aesthetics
photorealism
guidance-distillation
open-weights
2026年6月25日
Goku: Flow Based Video Generative Foundation Models
video-generation
text-to-image
image-to-video
rectified-flow
dit
joint-image-video
3d-vae
bytedance
2026年6月25日
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
text-to-image
dit
mmdit
moe
sparse-moe
flow-matching
distillation
dmd
gan
open-source
2026年6月25日
Imagen 4 / Imagen 4 Ultra / Imagen 4 Fast
text-to-image
diffusion
typography
imagen
google-deepmind
closed-source
synthid
2026年6月25日
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
video-generation
text-to-video
image-to-video
text-to-image
image-editing
latent-diffusion
flow-matching
dit
crossdit
nabla
sparse-attention
open-source
russian
2026年6月25日
ERNIE-Image Technical Report
text-to-image
dit
single-stream
latent-diffusion
flow-matching
dpo
distillation
dmd
text-rendering
open-weights
chinese