AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: reinforcement-learning
此标签下有50条笔记。
2026年7月16日
Habitat: A Platform for Embodied AI Research
embodied-ai
simulator
pointgoal-navigation
photorealistic
matterport3d
gibson
reinforcement-learning
benchmark
sim2real
2026年7月16日
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
benchmark
meta-rl
multi-task-rl
robotic-manipulation
mujoco
sawyer
reinforcement-learning
generalization
2026年7月16日
Learning Quadrupedal Locomotion over Challenging Terrain
legged-locomotion
quadruped
anymal
teacher-student
privileged-learning
tcn
sim-to-real
terrain-curriculum
proprioception
reinforcement-learning
2026年7月16日
Brax — A Differentiable Physics Engine for Large Scale Rigid Body Simulation
physics-simulation
differentiable-simulation
jax
xla
tpu
gpu
reinforcement-learning
rigid-body
maximal-coordinates
vectorized-envs
2026年7月16日
Habitat 2.0: Training Home Assistants to Rearrange their Habitat
embodied-ai
simulator
rearrangement
mobile-manipulation
replicacad
rigid-body-physics
reinforcement-learning
sense-plan-act
home-assistant-benchmark
ddppo
2026年7月16日
Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
physics-simulation
gpu-simulation
physx
reinforcement-learning
vectorized-envs
sim-to-real
dexterous-manipulation
legged-locomotion
tensor-api
ppo
2026年7月16日
ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
sim-infra
manipulation
benchmark
sapien
partnet-mobility
point-cloud
learning-from-demonstrations
reinforcement-learning
generalization
2026年7月16日
RMA: Rapid Motor Adaptation for Legged Robots
legged-locomotion
quadruped
unitree-a1
rma
online-adaptation
privileged-learning
teacher-student
system-identification
sim-to-real
proprioception
ppo
reinforcement-learning
2026年7月16日
Deep Whole-Body Control: Learning a Unified Policy for Manipulation and Locomotion
whole-body-control
legged-manipulator
quadruped-manipulation
reinforcement-learning
sim-to-real
advantage-mixing
regularized-online-adaptation
ppo
isaac-gym
unitree-go1
2026年7月16日
Barkour: Benchmarking Animal-level Agility with Quadruped Robots
quadruped
legged-locomotion
agility-benchmark
sim-to-real
reinforcement-learning
transformer-distillation
teacher-student
domain-randomization
2026年7月16日
DreamWaQ: Learning Robust Quadrupedal Locomotion with Implicit Terrain Imagination via Deep Reinforcement Learning
legged-locomotion
quadruped
reinforcement-learning
proprioception
asymmetric-actor-critic
terrain-imagination
variational-autoencoder
sim-to-real
unitree-a1
domain-randomization
2026年7月16日
Extreme Parkour with Legged Robots
legged-locomotion
quadruped
parkour
sim-to-real
egocentric-depth
teacher-student-distillation
reinforcement-learning
unitree-a1
inner-product-reward
roa
2026年7月16日
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
maniskill2
benchmark
sim-infra
sapien
mpm-soft-body
render-server
gpu-simulation
imitation-learning
reinforcement-learning
sim2real
2026年7月16日
MuJoCo XLA (MJX)
mjx
mujoco
jax
xla
gpu-simulation
tpu
differentiable-simulation
reinforcement-learning
physics-engine
warp
2026年7月16日
Real-World Humanoid Locomotion with Reinforcement Learning
humanoid
locomotion
reinforcement-learning
sim-to-real
causal-transformer
in-context-learning
domain-randomization
isaac-gym
digit
bair
2026年7月16日
RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation
robogen
generative-simulation
llm-agent
task-generation
reward-generation
objaverse
partnet-mobility
genesis
reinforcement-learning
motion-planning
2026年7月16日
Robot Parkour Learning
legged-locomotion
quadruped
parkour
sim-to-real
reinforcement-learning
privileged-distillation
depth-vision
unitree
2026年7月16日
DemoStart: Demonstration-led Auto-curriculum Applied to Sim-to-real with Multi-fingered Robots
dexterous-manipulation
sim-to-real
auto-curriculum
reinforcement-learning
zero-shot-transfer
multi-fingered-hand
behavior-cloning
domain-randomization
mpo
sparse-reward
2026年7月16日
DPPO: Diffusion Policy Policy Optimization
manipulation
diffusion-policy
reinforcement-learning
ppo
policy-gradient
rl-finetuning
sim-to-real
action-chunking
ddim
embodied-ai
2026年7月16日
Expressive Whole-Body Control for Humanoid Robots
humanoid
whole-body-control
unitree-h1
motion-imitation
reinforcement-learning
cmu-mocap
sim2real
rss-2024
ppo
motion-retargeting
2026年7月16日
Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation (H2O)
humanoid
teleoperation
whole-body-control
motion-imitation
sim-to-real
reinforcement-learning
retargeting
unitree-h1
amass
rgb-camera
2026年7月16日
Humanoid Parkour Learning
humanoid
parkour
legged-locomotion
sim-to-real
reinforcement-learning
privileged-distillation
depth-vision
unitree-h1
whole-body-control
2026年7月16日
Learning Multi-Modal Whole-Body Control for Real-World Humanoid Robots
humanoid
whole-body-control
digit-v3
motion-tracking
masked-conditioning
reinforcement-learning
multi-modal-control
sim2real
curriculum-learning
lstm
2026年7月16日
UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers
quadruped-manipulation
whole-body-control
umi
diffusion-policy
cross-embodiment
sim2real
reinforcement-learning
end-effector-trajectory
task-frame-tracking
corl-2024
2026年7月16日
WoCoCo: Learning Whole-Body Humanoid Control with Sequential Contacts
humanoid
multi-contact
whole-body-control
reinforcement-learning
curiosity-exploration
sim-to-real
unitree-h1
parkour
loco-manipulation
ppo
Exploration
2026年7月16日
ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills
humanoid
sim-to-real
residual-action
whole-body-control
motion-tracking
unitree-g1
reinforcement-learning
delta-action
dynamics-gap
2026年7月16日
The Developments and Challenges towards Dexterous and Embodied Robotic Manipulation: A Survey
survey
dexterous-manipulation
embodied-intelligence
multi-fingered-hand
teleoperation
imitation-learning
reinforcement-learning
data-collection
sim2real
human-to-robot-gap
2026年7月16日
A Survey on Efficient Vision-Language-Action Models
survey
vla
efficient-vla
model-compression
quantization
token-pruning
mixture-of-experts
action-tokenization
data-collection
reinforcement-learning
2026年7月16日
HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit
humanoid
loco-manipulation
teleoperation
exoskeleton
reinforcement-learning
whole-body-control
unitree-g1
fourier-gr1
imitation-learning
mocap-free
2026年7月16日
HuB: Learning Extreme Humanoid Balance
humanoid
balance-control
reinforcement-learning
sim-to-real
motion-retargeting
unitree-g1
teacher-student-distillation
whole-body-control
quasi-static-balance
2026年7月16日
HugWBC: A Unified and General Humanoid Whole-Body Controller for Versatile Locomotion
humanoid
whole-body-control
loco-manipulation
command-space
gait-control
teleoperation-intervention
sim-to-real
reinforcement-learning
unitree-h1
isaacgym
2026年7月16日
Learning Getting-Up Policies for Real-World Humanoid Robots
humanoid
fall-recovery
getting-up
whole-body-control
curriculum-learning
sim-to-real
reinforcement-learning
unitree-g1
isaacgym
ppo
2026年7月16日
KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills
humanoid
whole-body-control
motion-tracking
unitree-g1
adaptive-curriculum
bi-level-optimization
sim-to-real
reinforcement-learning
motion-retargeting
2026年7月16日
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
survey
vla
vision-language-action
robotic-manipulation
taxonomy
monolithic-models
hierarchical-models
world-models
reinforcement-learning
embodied-ai
2026年7月16日
MuJoCo Playground: An Open-Source Framework for GPU-Accelerated Robot Learning and Sim-to-Real Transfer
sim-infra
mjx
mujoco
gpu-simulation
sim-to-real
reinforcement-learning
madrona
batch-rendering
locomotion
manipulation
2026年7月16日
OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
humanoid
motion-retargeting
interaction-mesh
loco-manipulation
data-augmentation
sim-to-real
unitree-g1
trajectory-optimization
reinforcement-learning
2026年7月16日
1X NEO + Redwood AI (learned world-model policy)
humanoid
home-robot
vla
world-model
diffusion-policy
whole-body-control
inverse-dynamics-model
video-generation
reinforcement-learning
consumer-robotics
2026年7月16日
π*0.6: a VLA That Learns From Experience (RECAP)
vla
reinforcement-learning
offline-rl
advantage-conditioning
flow-matching
robot-learning
dagger
value-function
2026年7月16日
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
flow-matching
rectified-flow
shortcut-models
reinforcement-learning
ppo
policy-gradient
rl-finetuning
action-chunking
noise-injection
embodied-ai
2026年7月16日
ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
humanoid
loco-manipulation
residual-learning
motion-tracking
whole-body-control
sim-to-real
unitree-g1
reinforcement-learning
contact-reward
ppo
2026年7月16日
RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
simulation
metasim
cross-simulator
cross-embodiment
robot-learning-benchmark
imitation-learning
reinforcement-learning
world-model
sim-to-real
teleoperation
2026年7月16日
Survey of Vision-Language-Action Models for Embodied Manipulation
survey
vision-language-action
embodied-manipulation
robot-learning
imitation-learning
reinforcement-learning
action-tokenization
chain-of-thought
hierarchical-vla
benchmark
2026年7月16日
WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation Control
vla
humanoid
loco-manipulation
latent-action-model
vq-vae
reinforcement-learning
whole-body-control
agibot-x2
prismatic-vlm
iclr-2026
2026年7月16日
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
benchmark
reward-model
vision-language-model
reinforcement-learning
robot-learning
open-x-embodiment
roboarena
qwen3-vl
data-augmentation
2026年7月16日
Policy Improvement by Planning with Gumbel (Gumbel MuZero)
reinforcement-learning
mcts
alphazero
muzero
policy-improvement
gumbel-top-k
sequential-halving
tree-search
model-based-rl
2026年7月16日
LightZero: A Unified Benchmark for Monte Carlo Tree Search in General Sequential Decision Scenarios
world-model
mcts
muzero
alphazero
model-based-rl
benchmark
open-source-toolkit
tree-search
reinforcement-learning
2026年7月16日
ReZero: Boosting MCTS-based Algorithms by Backward-view and Entire-buffer Reanalyze
reinforcement-learning
mcts
muzero
efficientzero
reanalyze
sample-efficiency
tree-search
bandit-theory
lightzero
2026年7月16日
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
physical-ai
vlm
embodied-reasoning
chain-of-thought
grpo
reinforcement-learning
ontology
benchmark
cosmos
robotics
2026年7月16日
ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
world-model
3d-gaussian-splatting
panoramic-video
spatial-generative-primitive
diffusion-transformer
video-diffusion
reinforcement-learning
scene-reconstruction
alibaba
amap
2026年7月16日
DreamX-World 1.0: A General-Purpose Interactive World Model
world-model
video-generation
camera-control
autoregressive
diffusion-transformer
memory-conditioning
event-control
reinforcement-learning
wan2.2
alibaba