AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: ppo
此标签下有19条笔记。
2026年7月16日
Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
sim-to-real
quadruped-locomotion
minitaur
ppo
dynamics-randomization
actuator-modeling
reality-gap
deep-rl
pybullet
legged-robots
2026年7月16日
Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
physics-simulation
gpu-simulation
physx
reinforcement-learning
vectorized-envs
sim-to-real
dexterous-manipulation
legged-locomotion
tensor-api
ppo
2026年7月16日
Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning
legged-locomotion
quadruped
isaac-gym
ppo
sim-to-real
curriculum
gpu-rl
anymal
on-policy
massively-parallel
2026年7月16日
RMA: Rapid Motor Adaptation for Legged Robots
legged-locomotion
quadruped
unitree-a1
rma
online-adaptation
privileged-learning
teacher-student
system-identification
sim-to-real
proprioception
ppo
reinforcement-learning
2026年7月16日
Deep Whole-Body Control: Learning a Unified Policy for Manipulation and Locomotion
whole-body-control
legged-manipulator
quadruped-manipulation
reinforcement-learning
sim-to-real
advantage-mixing
regularized-online-adaptation
ppo
isaac-gym
unitree-go1
2026年7月16日
Walk These Ways: Tuning Robot Control for Generalization with Multiplicity of Behavior
quadruped-locomotion
multiplicity-of-behavior
gait-conditioned-rl
sim-to-real
ppo
isaac-gym
unitree-go1
domain-randomization
raibert-heuristic
out-of-distribution
2026年7月16日
DTC: Deep Tracking Control
legged-locomotion
quadruped
anymal
trajectory-optimization
hybrid-control
tamols
ppo
asymmetric-actor-critic
sim-to-real
sparse-terrain
2026年7月16日
Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion
legged-locomotion
quadruped
reach-avoid
hamilton-jacobi
safe-rl
collision-avoidance
dual-policy
ppo
sim-to-real
unitree-go1
2026年7月16日
DPPO: Diffusion Policy Policy Optimization
manipulation
diffusion-policy
reinforcement-learning
ppo
policy-gradient
rl-finetuning
sim-to-real
action-chunking
ddim
embodied-ai
2026年7月16日
Expressive Whole-Body Control for Humanoid Robots
humanoid
whole-body-control
unitree-h1
motion-imitation
reinforcement-learning
cmu-mocap
sim2real
rss-2024
ppo
motion-retargeting
2026年7月16日
HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots
humanoid
whole-body-control
unitree-h1
motion-imitation
policy-distillation
dagger
multi-mode
sim2real
ppo
icra-2025
2026年7月16日
Learning from Massive Human Videos for Universal Humanoid Pose Control
humanoid
whole-body-control
action-tokenization
vq-vae
transformer
motion-retargeting
internet-video
text-conditioned-control
unitree-h1
ppo
2026年7月16日
WoCoCo: Learning Whole-Body Humanoid Control with Sequential Contacts
humanoid
multi-contact
whole-body-control
reinforcement-learning
curiosity-exploration
sim-to-real
unitree-h1
parkour
loco-manipulation
ppo
Exploration
2026年7月16日
FALCON: Learning Force-Adaptive Humanoid Loco-Manipulation
humanoid
loco-manipulation
force-adaptation
dual-agent-rl
whole-body-control
unitree-g1
booster-t1
torque-limit-curriculum
sim-to-real
ppo
2026年7月16日
Learning Getting-Up Policies for Real-World Humanoid Robots
humanoid
fall-recovery
getting-up
whole-body-control
curriculum-learning
sim-to-real
reinforcement-learning
unitree-g1
isaacgym
ppo
2026年7月16日
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
flow-matching
rectified-flow
shortcut-models
reinforcement-learning
ppo
policy-gradient
rl-finetuning
action-chunking
noise-injection
embodied-ai
2026年7月16日
ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
humanoid
loco-manipulation
residual-learning
motion-tracking
whole-body-control
sim-to-real
unitree-g1
reinforcement-learning
contact-reward
ppo
2026年7月16日
Model-Based Reinforcement Learning for Atari
world-model
model-based-rl
video-prediction
atari
stochastic-discrete-latent
sample-efficiency
dyna
ppo
tensor2tensor
2026年6月25日
DDPO: Training Diffusion Models with Reinforcement Learning
diffusion
rl
rlhf
rlaif
policy-gradient
ppo
mdp
text-to-image
reward-optimization
vlm-feedback