AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: transformer
此标签下有34条笔记。
2026年7月16日
Behavior Transformers: Cloning k modes with one stone
behavior-cloning
imitation-learning
multimodal-action
action-discretization
transformer
minGPT
offline-learning
k-means
manipulation
2026年7月16日
Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions
offline-rl
q-learning
autoregressive-q-function
conservative-q-learning
action-discretization
transformer
rt-1-backbone
saycan
real-robot
corl-2023
2026年7月16日
ViNT: A Foundation Model for Visual Navigation
visual-navigation
foundation-model
image-goal
cross-embodiment
transformer
efficientnet
topological-graph
diffusion-subgoal
prompt-tuning
mobile-robot
2026年7月16日
BAKU: An Efficient Transformer for Multi-Task Policy Learning
multi-task-policy-learning
imitation-learning
transformer
action-chunking
film-conditioning
action-head
behavior-cloning
robot-learning
2026年7月16日
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
urban-navigation
point-goal-navigation
imitation-learning
visual-odometry
web-scale-video
dinov2
transformer
quadruped
data-scaling
cross-embodiment
2026年7月16日
HPT: Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
embodied-ai
robot-learning
cross-embodiment
policy-pretraining
transformer
scaling-laws
proprioception
behavior-cloning
2026年7月16日
RISE: 3D Perception Makes Real-World Robot Imitation Simple and Effective
manipulation
imitation-learning
point-cloud
sparse-convolution
minkowski-engine
diffusion-policy
transformer
real-world-robot
iros2024
2026年7月16日
Learning from Massive Human Videos for Universal Humanoid Pose Control
humanoid
whole-body-control
action-tokenization
vq-vae
transformer
motion-retargeting
internet-video
text-conditioned-control
unitree-h1
ppo
2026年7月16日
Transformers are Sample-Efficient World Models (IRIS)
world-model
model-based-rl
transformer
discrete-autoencoder
vqvae
autoregressive
image-tokens
atari-100k
sample-efficiency
latent-imagination
Superhuman
2026年7月16日
TransDreamer: Reinforcement Learning with Transformer World Models
world-model
model-based-rl
transformer
state-space-model
tssm
dreamer
memory-based-reasoning
actor-critic
partial-observability
2026年7月16日
STORM: Efficient Stochastic Transformer based World Models for Reinforcement Learning
world-model
model-based-rl
transformer
categorical-vae
atari-100k
imagination
actor-critic
sample-efficiency
2026年7月16日
Efficient World Models with Context-Aware Tokenization (Δ-IRIS)
world-model
model-based-rl
transformer
discrete-autoencoder
vqvae
autoregressive
delta-tokens
crafter
atari-100k
imagination
Superhuman
2026年7月16日
Point-JEPA: A Joint Embedding Predictive Architecture for Self-Supervised Learning on Point Cloud
jepa
self-supervised
point-cloud
3d-vision
latent-prediction
non-generative
masking
representation-learning
transformer
2026年7月16日
T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular Data
jepa
self-supervised
tabular-data
latent-prediction
non-contrastive
representation-learning
transformer
masking
regularization-token
2026年7月16日
UniZero: Generalized and Efficient Planning with Scalable Latent World Models
world-model
transformer
mcts
muzero
model-based-rl
atari-100k
dmcontrol
multitask-learning
lightzero
latent-world-model
2026年7月16日
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
world-model
video-generation
masked-token-prediction
vqgan
transformer
multimodal-conditioning
action-conditioning
autonomous-driving
parallel-decoding
2026年7月16日
World and Human Action Models towards gameplay ideation (WHAM / Muse)
world-model
game-wm
autoregressive
transformer
vqgan
tokenizer
action-conditioning
bleeding-edge
nature
muse
2026年7月16日
SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
world-model
autonomous-driving
4d-occupancy
sparse-query
trajectory-conditioning
transformer
feed-forward
nuscenes
occ3d
chamfer-distance
2026年7月16日
世界模型架构演进(RSSM · Transformer/AR · 扩散 · JEPA · 因果 tokenizer · 动作条件)
world-model
architecture
rssm
transformer
diffusion
jepa
tokenizer
action-conditioning
2026年6月25日
Taming Transformers for High-Resolution Image Synthesis (VQGAN)
vqgan
vq-vae
autoregressive
transformer
discrete-tokenizer
codebook
gan
perceptual-loss
image-synthesis
2026年6月25日
CogView: Mastering Text-to-Image Generation via Transformers
autoregressive
vqvae
transformer
gpt
text-to-image
chinese
fp16-stability
sandwich-ln
pb-relax
2026年6月25日
DALL·E: Zero-Shot Text-to-Image Generation
autoregressive
transformer
dvae
vqvae
t2i
zero-shot
sparse-attention
clip-rerank
2026年6月25日
VideoGPT: Video Generation using VQ-VAE and Transformers
video-generation
vq-vae
autoregressive
transformer
gpt
axial-attention
likelihood-based
2026年6月25日
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
text-to-video
autoregressive
transformer
vqvae
cogview2
open-source
2026年6月25日
MaskGIT: Masked Generative Image Transformer
masked-generation
parallel-decoding
vq
transformer
image-synthesis
non-autoregressive
params
steps
2026年6月25日
Point·E: A System for Generating 3D Point Clouds from Complex Prompts
text-to-3d
point-cloud
diffusion
transformer
glide
clip
image-to-3d
2026年6月25日
All are Worth Words: A ViT Backbone for Diffusion Models (U-ViT)
diffusion
vit
backbone
transformer
long-skip-connection
latent-diffusion
t2i
class-conditional
Layers
Heads
Params
2026年6月25日
UniDiffuser: One Transformer Fits All Distributions in Multi-Modal Diffusion
unified
multimodal
diffusion
transformer
u-vit
t2i
i2t
joint-generation
latent-diffusion
2026年6月25日
LRM: Large Reconstruction Model for Single Image to 3D
3d-reconstruction
single-image-to-3d
triplane
nerf
transformer
feed-forward
objaverse
2026年6月25日
Muse: Text-To-Image Generation via Masked Generative Transformers
t2i
masked-generative
transformer
vqgan
parallel-decoding
maskgit
t5
discrete-tokens
2026年6月25日
MusicGen / AudioCraft: Simple and Controllable Music Generation
music-generation
text-to-music
audio-lm
encodec
rvq
codebook-interleaving
autoregressive
transformer
open-source
2026年6月25日
W.A.L.T: Photorealistic Video Generation with Diffusion Models
video-diffusion
latent-diffusion
transformer
dit
window-attention
causal-3d-vae
text-to-video
magvit-v2
2026年6月25日
PixArt-δ: Fast and Controllable Image Generation with Latent Consistency Models
pixart
dit
lcm
consistency-distillation
controlnet
transformer
text-to-image
few-step
huawei
2026年6月25日
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
unified
multimodal
diffusion
next-token
transformer
late-fusion
chameleon
vae