AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
Home
❯
omni
❯
2024
文件夹: omni/2024
此文件夹下有69条笔记。
2026年6月25日
Allegro: Open the Black Box of Commercial-Level Video Generation Model
text-to-video
dit
video-vae
3d-rope
full-attention
open-weights
flow-free-diffusion
rhymes-ai
2026年6月25日
Aurora (Grok Image Generation)
autoregressive
mixture-of-experts
next-token
interleaved
multimodal
image-generation
image-editing
photorealism
closed-source
grok
2026年6月25日
Baichuan-Omni
omni
mllm
audio
video
image
speech
open-source
vita
siglip
whisper
conv-gmlp
2026年6月25日
Chameleon: Mixed-Modal Early-Fusion Foundation Models
unified
early-fusion
token-based
autoregressive
mixed-modal
vqgan
image-tokenizer
multimodal-llm
2026年6月25日
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
text-to-video
diffusion-transformer
3d-vae
expert-adaln
dit
open-source
i2v
2026年6月25日
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
t2i
relay-diffusion
cascaded-diffusion
latent-diffusion
distillation
recaption
unet
dit
cogview
2026年6月25日
DALL·E 3 系列产品化(ChatGPT 集成 / API)
t2i
latent-diffusion
recaptioning
prompt-following
chatgpt
closed-source
openai
safety
system-card
2026年6月25日
ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
t2i
diffusion
llm-text-encoder
adapter
prompt-following
dpg-bench
resampler
adaln
sd15
sdxl
training-free-unet
可训练参数
2026年6月25日
Emu3: Next-Token Prediction is All You Need
unified
autoregressive
next-token-prediction
discrete-token
vision-tokenizer
t2i
t2v
vlm
movqgan
dpo
2026年6月25日
FLUX1.1 [pro]
t2i
flux
rectified-flow
flow-matching
mmdit
closed-source
api
high-resolution
2026年6月25日
FLUX.1 Tools (Fill / Canny / Depth / Redux)
flux
inpainting
outpainting
controlnet
structural-conditioning
image-variation
mmdit
flow-matching
guidance-distillation
redux
siglip
2026年6月25日
FLUX.1 suite (pro / dev / schnell)
t2i
rectified-flow
flow-matching
mmdit
dit
ladd
guidance-distillation
open-weights
bfl
2026年6月25日
Gemini 2.0 Flash 原生图像生成(Native Image Output)
native-image-output
interleaved-text-image
unified
multimodal-llm
conversational-editing
synthid
closed-source
nano-banana-lineage
2026年6月25日
Runway Gen-3 Alpha
text-to-video
image-to-video
video-generation
closed-source
world-model
diffusion
c2pa
2026年6月25日
Genie: Generative Interactive Environments
world-model
video
latent-action
unsupervised
maskgit
st-transformer
vq-vae
foundation-model
playable
agents
Params
2026年6月25日
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
speech-lm
spoken-chatbot
speech-tokenizer
flow-matching
interleaved-pretraining
streaming
end-to-end
glm
2026年6月25日
GPT-4o 原生图像生成 (4o image generation / gpt-image-1)
omni
autoregressive
native-image
image-generation
text-rendering
image-editing
multimodal
closed-source
2026年6月25日
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
autoregressive
visual-tokenizer
residual-diffusion
var
t2i
efficient-inference
1024px
Params
Step
2026年6月25日
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
t2i
diffusion-transformer
dit
chinese
bilingual
multi-resolution
rope
recaptioning
mllm
open-source
2026年6月25日
HunyuanVideo: A Systematic Framework For Large Video Generative Models
video
t2v
dit
flow-matching
mmdit
3d-vae
open-source
scaling-law
2026年6月25日
Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
3d-generation
image-to-3d
text-to-3d
multi-view-diffusion
sparse-view-reconstruction
triplane
sdf
feed-forward
2026年6月25日
Ideogram 2.0
t2i
typography
text-rendering
closed-source
diffusion
product-launch
2026年6月25日
Imagen 3
t2i
latent-diffusion
google
gemini
closed-source
synthetic-caption
human-eval
synthid
2026年6月25日
In-Context LoRA for Diffusion Transformers (IC-LoRA)
in-context
lora
dit
flux
image-set
task-agnostic
editing
customization
2026年6月25日
Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
t2i
autoregressive
var
next-scale-prediction
bitwise-tokenizer
bsq
infinite-vocabulary
self-correction
scaling-law
2026年6月25日
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
unified
autoregressive
vqgan
siglip
decoupled-encoder
any-to-any
deepseek
2026年6月25日
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
unified
multimodal
rectified-flow
autoregression
llm
t2i
vlm
deepseek
2026年6月25日
Kling (可灵) 视频生成大模型
video-generation
t2v
i2v
dit
3d-vae
spatiotemporal-attention
closed-source
kuaishou
2026年6月25日
Kolors(可图): Effective Training of Diffusion Model for Photorealistic Text-to-Image Synthesis
t2i
latent-diffusion
unet
sdxl
chatglm
bilingual
chinese-text-rendering
recaption
open-source
2026年6月25日
Liquid: Language Models are Scalable and Unified Multi-modal Generators
unified-multimodal
autoregressive
vqgan
discrete-token
next-token-prediction
scaling-law
text-to-image
t2i
mllm
2026年6月25日
LlamaGen: Autoregressive Model Beats Diffusion — Llama for Scalable Image Generation
autoregressive
next-token
image-tokenizer
vqgan
llama
t2i
class-conditional
vllm
2026年6月25日
LTX-Video: Realtime Video Latent Diffusion
video
t2v
i2v
dit
latent-diffusion
rectified-flow
video-vae
realtime
open-source
2026年6月25日
Lumiere: A Space-Time Diffusion Model for Video Generation
text-to-video
diffusion
space-time-unet
t2i-inflation
pixel-diffusion
image-to-video
video-inpainting
2026年6月25日
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
text-to-image
diffusion-transformer
next-dit
flow-matching
rectified-flow
rope
resolution-extrapolation
few-step-sampling
multilingual
unified-generation
2026年6月25日
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
dit
flow-matching
flag-dit
text-to-image
text-to-video
text-to-3d
text-to-speech
rope
resolution-extrapolation
unified-generation
2026年6月25日
Autoregressive Image Generation without Vector Quantization (MAR / Diffusion Loss)
autoregressive
continuous-token
diffusion-loss
masked-generation
mar
tokenizer-free
imagenet
kaiming-he
2026年6月25日
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
unified
mllm
instruction-tuning
vpit
continuous-visual-tokens
autoregressive
diffusion-autoencoder
llama-3
2026年6月25日
Midjourney V6 / V6.1
t2i
diffusion
aesthetics
closed-source
text-rendering
midjourney
2026年6月25日
Mochi 1 (preview)
text-to-video
diffusion-transformer
asymmdit
mmdit
3d-attention
flow-matching
open-source
apache-2.0
video-vae
2026年6月25日
Movie Gen: A Cast of Media Foundation Models
video-generation
text-to-video
flow-matching
video-editing
personalization
video-to-audio
llama3-backbone
temporal-autoencoder
2026年6月25日
OminiControl: Minimal and Universal Control for Diffusion Transformer
dit
flux
controlnet
subject-driven
image-conditioning
lora
rope
dataset
2026年6月25日
OmniGen: Unified Image Generation
unified-generation
diffusion
rectified-flow
image-editing
subject-driven
controlnet-free
instruction-following
x2i
phi-3
2026年6月25日
Open-Sora Plan / Open-Sora(2024 开源复现 Sora)
video
t2v
i2v
dit
diffusion
sora-reproduction
open-source
wf-vae
skiparse-attention
rectified-flow
2026年6月25日
Sora(Sora Turbo)公开发布
video-generation
text-to-video
diffusion-transformer
dit
spacetime-patches
world-simulator
recaptioning
closed-source
2026年6月25日
PixArt-δ: Fast and Controllable Image Generation with Latent Consistency Models
pixart
dit
lcm
consistency-distillation
controlnet
transformer
text-to-image
few-step
huawei
2026年6月25日
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
t2i
dit
diffusion-transformer
4k
kv-compression
weak-to-strong
efficient
pixart
Params
2026年6月25日
Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models
t2i
diffusion
dit
deep-fusion
llm-text-encoder
llama3
graphic-design
text-rendering
capsbench
closed-source
2026年6月25日
Pyramid Flow:金字塔式流匹配的高效视频生成
video-generation
flow-matching
pyramidal-flow
autoregressive
dit
mm-dit
efficient-training
2026年6月25日
Qwen2-Audio
audio-language-model
lalm
speech
asr
s2tt
voice-chat
dpo
whisper
qwen
2026年6月25日
Recraft V3 (red_panda)
t2i
closed-source
text-rendering
layout-control
vector-graphics
controlnet
design
api
ocr
2026年6月25日
REPA: Representation Alignment for Generation — Training Diffusion Transformers Is Easier Than You Think
diffusion-transformer
representation-alignment
dinov2
sit
dit
self-supervised
training-efficiency
flow-matching
imagenet
2026年6月25日
Sana: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
t2i
diffusion-transformer
linear-attention
deep-compression-ae
rectified-flow
efficient
on-device
gemma
2026年6月25日
SD3.5 Large Turbo 与 (Latent) Adversarial Diffusion Distillation
distillation
adversarial
few-step
mmdit
rectified-flow
turbo
gan
text-to-image
2026年6月25日
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
tts
speech-synthesis
zero-shot
voice-cloning
autoregressive
diffusion
dit
rl
voice-conversion
in-context-learning
2026年6月25日
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
unified
mllm
external-diffuser
sdxl
image-editing
llama2
vit-bridge
comprehension-generation
2026年6月25日
SeedEdit: Align Image Re-Generation to Image Editing
image-editing
instruction-editing
diffusion
t2i-bootstrap
self-distillation
causal-attention
iterative-alignment
bytedance-seed
2026年6月25日
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
unified
understanding-generation
discrete-diffusion
maskgit
autoregressive
magvit-v2
phi-1.5
2026年6月25日
SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
diffusion
flow-matching
stochastic-interpolant
dit
imagenet
t2i-backbone
sde
ode
2026年6月25日
Sora: Video generation models as world simulators
video-generation
text-to-video
diffusion-transformer
dit
spacetime-patches
world-simulator
recaptioning
native-resolution
scaling
closed-source
2026年6月25日
Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
audio
music-generation
latent-diffusion
text-to-audio
timing-conditioning
stereo
vae
clap
2026年6月25日
Stable Cascade (Würstchen v3)
text-to-image
latent-diffusion
cascade
wuerstchen
efficient
convnext
vqgan
semantic-compressor
open-weights
2026年6月25日
Stable Diffusion 3.5 (Large / Large Turbo / Medium)
t2i
mmdit
rectified-flow
open-weights
qk-norm
distillation
sd3
2026年6月25日
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (Stable Diffusion 3 / MMDiT)
t2i
rectified-flow
flow-matching
mmdit
dit
multimodal-transformer
dpo
scaling
open-weights
stability-ai
2026年6月25日
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
unified
multimodal
diffusion
next-token
transformer
late-fusion
chameleon
vae
2026年6月25日
TRELLIS: Structured 3D Latents for Scalable and Versatile 3D Generation
3d-generation
slat
rectified-flow
dit
sparse-voxel
dinov2
image-to-3d
text-to-3d
gaussian-splatting
mesh
2026年6月25日
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
image-editing
instruction-edit
dataset
region-based
diffusion
sdxl-turbo
prompt-to-prompt
neurips-2024
2026年6月25日
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction (VAR)
autoregressive
next-scale-prediction
image-generation
scaling-laws
vqvae
imagenet
neurips-best-paper
参数
Step
2026年6月25日
Veo 2
video-generation
text-to-video
image-to-video
closed-source
cinematography
latent-diffusion
synthid
4k
2026年6月25日
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
unified
autoregressive
next-token
visual-tokenizer
rq-vae
clip-alignment
vlm
image-generation
video-generation