AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: vqgan
此标签下有18条笔记。
2026年7月16日
World Model on Million-Length Video And Language With Blockwise RingAttention
world-model
ring-attention
long-context
million-token-context
video-language-model
autoregressive-transformer
vqgan
blockwise-parallel-transformer
tpu-training
rope-scaling
2026年7月16日
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
world-model
video-generation
masked-token-prediction
vqgan
transformer
multimodal-conditioning
action-conditioning
autonomous-driving
parallel-decoding
2026年7月16日
World and Human Action Models towards gameplay ideation (WHAM / Muse)
world-model
game-wm
autoregressive
transformer
vqgan
tokenizer
action-conditioning
bleeding-edge
nature
muse
2026年6月25日
Taming Transformers for High-Resolution Image Synthesis (VQGAN)
vqgan
vq-vae
autoregressive
transformer
discrete-tokenizer
codebook
gan
perceptual-loss
image-synthesis
2026年6月25日
ERNIE-ViLG: Unified Generative Pre-training for Bidirectional Vision-Language Generation
text-to-image
image-captioning
autoregressive
vqgan
unified
chinese
baidu
2026年6月25日
M6: A Chinese Multimodal Pretrainer
multimodal
chinese
pretraining
moe
text-to-image
vqgan
encoder-decoder
mixture-of-experts
2026年6月25日
Vector-quantized Image Modeling with Improved VQGAN (ViT-VQGAN)
tokenizer
vqgan
vit
vector-quantization
autoregressive
codebook
image-generation
representation-learning
2026年6月25日
VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance
vqgan
clip
text-to-image
clip-guidance
training-free
image-editing
ai-art
latent-optimization
eleutherai
2026年6月25日
Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors
t2i
autoregressive
vqgan
scene-control
segmentation
classifier-free-guidance
human-prior
meta
2026年6月25日
ModelScopeT2V(ModelScope Text-to-Video Technical Report)
text-to-video
latent-diffusion
spatio-temporal
unet3d
open-source
webvid
vqgan
2026年6月25日
Muse: Text-To-Image Generation via Masked Generative Transformers
t2i
masked-generative
transformer
vqgan
parallel-decoding
maskgit
t5
discrete-tokens
2026年6月25日
Würstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion Models
t2i
latent-diffusion
cascade
compression
vqgan
efficient-training
convnext
stable-cascade
2026年6月25日
Chameleon: Mixed-Modal Early-Fusion Foundation Models
unified
early-fusion
token-based
autoregressive
mixed-modal
vqgan
image-tokenizer
multimodal-llm
2026年6月25日
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
unified
autoregressive
vqgan
siglip
decoupled-encoder
any-to-any
deepseek
2026年6月25日
Liquid: Language Models are Scalable and Unified Multi-modal Generators
unified-multimodal
autoregressive
vqgan
discrete-token
next-token-prediction
scaling-law
text-to-image
t2i
mllm
2026年6月25日
LlamaGen: Autoregressive Model Beats Diffusion — Llama for Scalable Image Generation
autoregressive
next-token
image-tokenizer
vqgan
llama
t2i
class-conditional
vllm
2026年6月25日
Stable Cascade (Würstchen v3)
text-to-image
latent-diffusion
cascade
wuerstchen
efficient
convnext
vqgan
semantic-compressor
open-weights
2026年6月25日
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
autoregressive
t2i
unified-generation
decoder-only
vqgan
image-editing
controllable-generation
speculative-jacobi