AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: autoregressive
此标签下有74条笔记。
2026年7月16日
Galaxea G0.5 Technical Report
vla
autoregressive
action-tokenizer
chain-of-thought
cross-embodiment
visual-memory
qwen3.5
residual-vq
r1-lite
r1-pro
2026年7月16日
Transformers are Sample-Efficient World Models (IRIS)
world-model
model-based-rl
transformer
discrete-autoencoder
vqvae
autoregressive
image-tokens
atari-100k
sample-efficiency
latent-imagination
Superhuman
2026年7月16日
OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
world-model
autonomous-driving
3d-occupancy
vqvae
gpt
autoregressive
4d-forecasting
planning
nuscenes
occ3d
2026年7月16日
Transformer-based World Models Are Happy With 100k Interactions (TWM)
world-model
model-based-rl
transformer-xl
atari-100k
sample-efficiency
latent-imagination
autoregressive
actor-critic
kl-balancing
2026年7月16日
Genie 2: A Large-Scale Foundation World Model
world-model
foundation-world-model
latent-diffusion
autoregressive
action-conditioning
playable
embodied-agents
imagen-3
sima
game-wm
2026年7月16日
Efficient World Models with Context-Aware Tokenization (Δ-IRIS)
world-model
model-based-rl
transformer
discrete-autoencoder
vqvae
autoregressive
delta-tokens
crafter
atari-100k
imagination
Superhuman
2026年7月16日
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
video-tokenizer
autoregressive
holistic-tokenization
stochastic-vector-quantization
world-model
FVD
UCF-101
kinetics-600
ICLR2025
Tokens
2026年7月16日
MirageLSD: The First Live-Stream Diffusion AI Video Model
world-model
live-stream-diffusion
real-time-video
diffusion-forcing
autoregressive
zero-latency
video-transformation
mega-kernel
shortcut-distillation
2026年7月16日
Genie 3: A new frontier for world models
world-model
interactive
real-time
autoregressive
video-generation
embodied-agent
agi
promptable-world-events
sima
2026年7月16日
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
world-model
generative-latent-prediction
glp
autoregressive
video-diffusion
causal-swin-dpm
qwen2.5-vl
wan2.1
long-horizon
simulative-reasoning
2026年7月16日
World and Human Action Models towards gameplay ideation (WHAM / Muse)
world-model
game-wm
autoregressive
transformer
vqgan
tokenizer
action-conditioning
bleeding-edge
nature
muse
2026年7月16日
Odyssey-1: A Playable World Model
world-model
interactive-video
real-time-generation
action-conditioned
autoregressive
gaussian-splatting
odyssey
ex-wayve
generative-simulation
2026年7月16日
Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
world-model
game-generation
interactive-video
camera-control
action-conditioning
autoregressive
diffusion-transformer
distillation
aaa-games
open-source
2026年7月16日
DreamX-World 1.0: A General-Purpose Interactive World Model
world-model
video-generation
camera-control
autoregressive
diffusion-transformer
memory-conditioning
event-control
reinforcement-learning
wan2.2
alibaba
2026年7月16日
世界模型训练方法(想象力 RL · 扩散/流 · next-token · JEPA 回归 · 蒸馏)
world-model
training
imagination-rl
dreamer
muzero
diffusion
flow-matching
diffusion-forcing
autoregressive
jepa
distillation
2026年6月25日
Omni / 多模态生成技术演进调研 (2020 → 2026-07) · 主汇总
omni
multimodal-generation
text-to-image
image-editing
unified-understanding-generation
any-to-any
video-generation
diffusion
autoregressive
survey
2026年6月25日
Generative Pretraining from Pixels (Image GPT / iGPT)
autoregressive
pixel-transformer
generative-pretraining
representation-learning
gpt
unsupervised
2026年6月25日
Jukebox: A Generative Model for Music
music-generation
raw-audio
vq-vae
sparse-transformer
autoregressive
codec-tokens
lyrics-conditioning
2026年6月25日
Taming Transformers for High-Resolution Image Synthesis (VQGAN)
vqgan
vq-vae
autoregressive
transformer
discrete-tokenizer
codebook
gan
perceptual-loss
image-synthesis
2026年6月25日
CogView: Mastering Text-to-Image Generation via Transformers
autoregressive
vqvae
transformer
gpt
text-to-image
chinese
fp16-stability
sandwich-ln
pb-relax
2026年6月25日
DALL·E: Zero-Shot Text-to-Image Generation
autoregressive
transformer
dvae
vqvae
t2i
zero-shot
sparse-attention
clip-rerank
2026年6月25日
ERNIE-ViLG: Unified Generative Pre-training for Bidirectional Vision-Language Generation
text-to-image
image-captioning
autoregressive
vqgan
unified
chinese
baidu
2026年6月25日
GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
text-to-video
autoregressive
vq-vae
sparse-attention
t2v
howto100m
msr-vtt
2026年6月25日
NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion
unified-generation
autoregressive
vq-gan
3d-transformer
sparse-attention
text-to-image
text-to-video
image-editing
any-to-vision
2026年6月25日
VideoGPT: Video Generation using VQ-VAE and Transformers
video-generation
vq-vae
autoregressive
transformer
gpt
axial-attention
likelihood-based
2026年6月25日
Vector-quantized Image Modeling with Improved VQGAN (ViT-VQGAN)
tokenizer
vqgan
vit
vector-quantization
autoregressive
codebook
image-generation
representation-learning
2026年6月25日
AudioLM: a Language Modeling Approach to Audio Generation
audio
speech
music
neural-codec
semantic-tokens
autoregressive
soundstream
w2v-bert
textless-nlp
2026年6月25日
CM3: A Causal Masked Multimodal Model of the Internet
autoregressive
decoder-only
causal-masking
multimodal
html
vqvae-gan
zero-shot
infilling
entity-linking
unified
2026年6月25日
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
text-to-video
autoregressive
transformer
vqvae
cogview2
open-source
2026年6月25日
CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers
autoregressive
hierarchical-transformer
bilingual
vqvae
super-resolution
masked-generation
lopar
coglm
2026年6月25日
Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors
t2i
autoregressive
vqgan
scene-control
segmentation
classifier-free-guidance
human-prior
meta
2026年6月25日
Parti: Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
autoregressive
text-to-image
vit-vqgan
seq2seq
scaling
classifier-free-guidance
tpu
partiprompts
2026年6月25日
Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions
text-to-video
video-generation
c-vivit
maskgit
masked-transformer
vq
story-generation
autoregressive
2026年6月25日
CM3Leon: Scaling Autoregressive Multi-Modal Models (Pretraining and Instruction Tuning)
autoregressive
token-based
retrieval-augmented
text-to-image
image-to-text
instruction-tuning
contrastive-decoding
cm3
chameleon-lineage
2026年6月25日
Emu: Generative Pretraining in Multimodality
unified
autoregressive
lmm
multimodal-generation
interleaved
video-text
eva-clip
llama
stable-diffusion
in-context-learning
2026年6月25日
MusicGen / AudioCraft: Simple and Controllable Music Generation
music-generation
text-to-music
audio-lm
encodec
rvq
codebook-interleaving
autoregressive
transformer
open-source
2026年6月25日
MusicLM: Generating Music From Text
text-to-music
audio-generation
discrete-tokens
autoregressive
audiolm
soundstream
mulan
hierarchical
2026年6月25日
SEED: Planting a SEED of Vision in Large Language Model
unified
visual-tokenizer
discrete-tokens
vq
multimodal-llm
autoregressive
q-former
image-to-text
text-to-image
2026年6月25日
VALL-E: Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
tts
zero-shot
voice-cloning
neural-codec
audio-lm
in-context-learning
encodec
autoregressive
2026年6月25日
VideoPoet: A Large Language Model for Zero-Shot Video Generation
video
autoregressive
llm
discrete-token
multimodal
text-to-video
image-to-video
audio
magvit-v2
zero-shot
2026年6月25日
Aurora (Grok Image Generation)
autoregressive
mixture-of-experts
next-token
interleaved
multimodal
image-generation
image-editing
photorealism
closed-source
grok
2026年6月25日
Chameleon: Mixed-Modal Early-Fusion Foundation Models
unified
early-fusion
token-based
autoregressive
mixed-modal
vqgan
image-tokenizer
multimodal-llm
2026年6月25日
Emu3: Next-Token Prediction is All You Need
unified
autoregressive
next-token-prediction
discrete-token
vision-tokenizer
t2i
t2v
vlm
movqgan
dpo
2026年6月25日
GPT-4o 原生图像生成 (4o image generation / gpt-image-1)
omni
autoregressive
native-image
image-generation
text-rendering
image-editing
multimodal
closed-source
2026年6月25日
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
autoregressive
visual-tokenizer
residual-diffusion
var
t2i
efficient-inference
1024px
Params
Step
2026年6月25日
Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
t2i
autoregressive
var
next-scale-prediction
bitwise-tokenizer
bsq
infinite-vocabulary
self-correction
scaling-law
2026年6月25日
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
unified
autoregressive
vqgan
siglip
decoupled-encoder
any-to-any
deepseek
2026年6月25日
Liquid: Language Models are Scalable and Unified Multi-modal Generators
unified-multimodal
autoregressive
vqgan
discrete-token
next-token-prediction
scaling-law
text-to-image
t2i
mllm
2026年6月25日
LlamaGen: Autoregressive Model Beats Diffusion — Llama for Scalable Image Generation
autoregressive
next-token
image-tokenizer
vqgan
llama
t2i
class-conditional
vllm
2026年6月25日
Autoregressive Image Generation without Vector Quantization (MAR / Diffusion Loss)
autoregressive
continuous-token
diffusion-loss
masked-generation
mar
tokenizer-free
imagenet
kaiming-he
2026年6月25日
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
unified
mllm
instruction-tuning
vpit
continuous-visual-tokens
autoregressive
diffusion-autoencoder
llama-3
2026年6月25日
Pyramid Flow:金字塔式流匹配的高效视频生成
video-generation
flow-matching
pyramidal-flow
autoregressive
dit
mm-dit
efficient-training
2026年6月25日
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
tts
speech-synthesis
zero-shot
voice-cloning
autoregressive
diffusion
dit
rl
voice-conversion
in-context-learning
2026年6月25日
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
unified
understanding-generation
discrete-diffusion
maskgit
autoregressive
magvit-v2
phi-1.5
2026年6月25日
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction (VAR)
autoregressive
next-scale-prediction
image-generation
scaling-laws
vqvae
imagenet
neurips-best-paper
参数
Step
2026年6月25日
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
unified
autoregressive
next-token
visual-tokenizer
rq-vae
clip-alignment
vlm
image-generation
video-generation
2026年6月25日
Cosmos World Foundation Model Platform for Physical AI (Cosmos-Predict1)
world-model
physical-ai
video-generation
diffusion
autoregressive
dit
tokenizer
robotics
autonomous-driving
open-weight
2026年6月25日
Emu3.5: Native Multimodal Models are World Learners
native-multimodal
autoregressive
next-token-prediction
world-model
interleaved
x2i
discrete-diffusion
open-source
2026年6月25日
Fractal Generative Models (FractalGen)
fractal
autoregressive
mar
pixel-by-pixel
image-generation
imagenet
divide-and-conquer
likelihood
params
2026年6月25日
Imagen 4 Fast / Imagen 4(Ultra) GA 与 GPT Image 1 mini
t2i
closed-source
api
cost-tier
imagen-4
gpt-image-1
autoregressive
latent-diffusion
synthid
c2pa
2026年6月25日
GPT Image 1 / 4o 原生图像生成(GPT-4o native image generation)
autoregressive
omnimodal
text-rendering
image-editing
instruction-following
c2pa
gpt-4o
closed-source
2026年6月25日
HunyuanImage 3.0 Technical Report
t2i
unified-multimodal
autoregressive
moe
diffusion
transfusion
chain-of-thought
open-source
rlhf
2026年6月25日
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
unified-multimodal
autoregressive
decoupled-visual-encoding
t2i
vlm
open-source
deepseek
2026年6月25日
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
autoregressive
t2i
unified-generation
decoder-only
vqgan
image-editing
controllable-generation
speculative-jacobi
2026年6月25日
MAGI-1: Autoregressive Video Generation at Scale
video
autoregressive
diffusion
flow-matching
chunk-wise
world-model
streaming
open-weights
2026年6月25日
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
autoregressive
continuous-tokens
flow-matching
next-token-prediction
t2i
image-editing
qwen2.5
flux-vae
nextstep-grpo
flowgrpo
2026年6月25日
Show-o2: Improved Native Unified Multimodal Models
unified
understanding-generation
autoregressive
flow-matching
3d-causal-vae
video
qwen2.5
omni-attention
2026年6月25日
SkyReels-V2: Infinite-Length Film Generative Model
video
diffusion-forcing
autoregressive
flow-matching
dit
dpo
long-video
open-source
2026年6月25日
X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
unified
autoregressive
discrete-token
rl
grpo
text-rendering
image-generation
image-understanding
tokenizer
siglip
flux
2026年6月25日
GPT Image 2 / ChatGPT Images 2.0(含 Thinking mode)
native-image-gen
autoregressive
thinking-mode
agentic
image-editing
text-rendering
4k
c2pa
watermark
closed-source
2026年6月25日
OmniGen-AR: AutoRegressive Any-to-Image Generation
autoregressive
any-to-image
unified-generation
next-token
visual-tokenizer
disentangled-causal-attention
image-editing
text-to-video
neurips-2025
2026年6月25日
统一/Omni 模型族横向对比:Chameleon · Emu · Janus · BAGEL · OmniGen · Show-o · VAR 谱系
unified-multimodal
omni
autoregressive
diffusion
flow-matching
next-scale-prediction
understanding-generation
image-editing
deep-dive
2026年6月25日
模型架构演进:从 U-Net 扩散到统一 omni 骨干(2020–2026)
omni
architecture
diffusion
dit
mmdit
rectified-flow
autoregressive
visual-tokenizer
vae
text-encoder
unified-multimodal
survey
2026年6月25日
统一理解生成 & any-to-any 全模态专题
unified
understanding-generation
any-to-any
omni
early-fusion
decoupled-encoder
diffusion-ar-hybrid
autoregressive
rectified-flow
thinker-talker