全量来源索引(按发布年月)
去重后 294 条一手来源(按
date年份归组;去重键为原始 URL)。
2017(2 条)
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World — OpenAI / UC Berkeley · 2017-03 · paper · [sim-infra] — https://arxiv.org/abs/1703.06907
- AI2-THOR: An Interactive 3D Environment for Visual AI — Allen Institute for AI (AI2) · 2017-12 · paper · [sim-infra] — https://arxiv.org/abs/1712.05474
2018(2 条)
- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots — Google Brain, X, Google DeepMind · 2018-04 · paper · [locomotion-nav] — https://arxiv.org/abs/1804.10332
- RoboTurk: A Crowdsourcing Platform for Robotic Skill Learning through Imitation — Stanford University · 2018-11 · paper · [data] — https://arxiv.org/abs/1811.02790
2019(5 条)
- Habitat: A Platform for Embodied AI Research — Meta FAIR / Facebook AI · 2019-04 · paper · [sim-infra] — https://arxiv.org/abs/1904.01201
- RLBench: The Robot Learning Benchmark & Learning Environment — Imperial College London (Dyson Robotics Lab) · 2019-09 · paper · [benchmark] — https://arxiv.org/abs/1909.12271
- DiffTaichi: Differentiable Programming for Physical Simulation — MIT · 2019-10 · paper · [sim-infra] — https://arxiv.org/abs/1910.00935
- Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning — UC Berkeley, Stanford, Columbia, USC, Google · 2019-10 · paper · [benchmark] — https://arxiv.org/abs/1910.10897
- RoboNet: Large-Scale Multi-Robot Learning — UC Berkeley, Stanford, University of Pennsylvania, CMU · 2019-10 · paper · [data] — https://arxiv.org/abs/1910.11215
2020(4 条)
- SAPIEN: A SimulAted Part-based Interactive ENvironment — UC San Diego / Stanford / SFU (Hao Su Lab) · 2020-03 · paper · [sim-infra] — https://arxiv.org/abs/2003.08515
- ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation — MIT-IBM Watson AI Lab / MIT / Harvard University / Stanford University · 2020-07 · paper · [sim-infra] — https://arxiv.org/abs/2007.04954
- LaND: Learning to Navigate from Disengagements — UC Berkeley (BAIR) · 2020-10 · paper · [locomotion-nav] — https://arxiv.org/abs/2010.04689
- Learning Quadrupedal Locomotion over Challenging Terrain — ETH Zurich & Intel · 2020-10 · paper · [locomotion-nav] — https://arxiv.org/abs/2010.11251
2021(13 条)
- Brax — A Differentiable Physics Engine for Large Scale Rigid Body Simulation — Google Research · 2021-06 · paper · [sim-infra] — https://arxiv.org/abs/2106.13281
- Habitat 2.0: Training Home Assistants to Rearrange their Habitat — Meta FAIR · 2021-06 · paper · [sim-infra] — https://arxiv.org/abs/2106.14405
- ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations — UC San Diego (Hao Su Lab) · 2021-07 · paper · [sim-infra] — https://arxiv.org/abs/2107.14483
- RMA: Rapid Motor Adaptation for Legged Robots — UC Berkeley, CMU, Facebook AI Research · 2021-07 · paper · [locomotion-nav] — https://arxiv.org/abs/2107.04034
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning — NVIDIA · 2021-08 · paper · [sim-infra] — https://arxiv.org/abs/2108.10470
- iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks — Stanford University (Stanford Vision & Learning Lab / SVL) · 2021-08 · paper · [sim-infra] — https://arxiv.org/abs/2108.03272
- BridgeData: Boosting Generalization of Robotic Skills with Cross-Domain Datasets — UC Berkeley · 2021-09 · paper · [data] — https://arxiv.org/abs/2109.13396
- Implicit Behavioral Cloning (IBC) — Robotics at Google · 2021-09 · paper · [manipulation] — https://arxiv.org/abs/2109.00137
- Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning — ETH Zurich (Robotic Systems Lab), NVIDIA · 2021-09 · paper · [locomotion-nav] — https://arxiv.org/abs/2109.11978
- Ego4D: Around the World in 3,000 Hours of Egocentric Video — Meta AI (FAIR) + consortium · 2021-10 · paper · [data] — https://arxiv.org/abs/2110.07058
- MuJoCo: A Physics Engine for Model-Based Control — DeepMind / Roboti LLC (E. Todorov) · 2021-10 · tech-report · [sim-infra] — https://mujoco.org/
- NVIDIA Isaac Sim — NVIDIA · 2021-11 · blog · [sim-infra] — https://developer.nvidia.com/isaac/sim
- CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks — University of Freiburg · 2021-12 · paper · [benchmark] — https://arxiv.org/abs/2112.03227
2022(17 条)
- Learning Robust Perceptive Locomotion for Quadrupedal Robots in the Wild — ETH Zurich · 2022-01 · paper · [locomotion-nav] — https://arxiv.org/abs/2201.08117
- BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning — Google (Robotics at Google) / Everyday Robots / UC Berkeley / Stanford · 2022-02 · paper · [manipulation] — https://arxiv.org/abs/2202.02005
- NVIDIA Warp: A High-performance Python Framework for GPU Simulation and Graphics — NVIDIA · 2022-03 · tech-report · [sim-infra] — https://github.com/NVIDIA/warp
- R3M: A Universal Visual Representation for Robot Manipulation — Stanford University / Meta AI · 2022-03 · paper · [data] — https://arxiv.org/abs/2203.12601
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan) — Google (Robotics at Google / Everyday Robots) · 2022-04 · paper · [vla] — https://arxiv.org/abs/2204.01691
- Behavior Transformers: Cloning k modes with one stone — New York University · 2022-06 · paper · [manipulation] — https://arxiv.org/abs/2206.11251
- ProcTHOR: Large-Scale Embodied AI Using Procedural Generation — Allen Institute for AI (AI2) · 2022-06 · paper · [sim-infra] — https://arxiv.org/abs/2206.06994
- Inner Monologue: Embodied Reasoning through Planning with Language Models — Google (Robotics at Google) · 2022-07 · paper · [vla] — https://arxiv.org/abs/2207.05608
- Code as Policies: Language Model Programs for Embodied Control — Google (Robotics at Google) · 2022-09 · paper · [vla] — https://arxiv.org/abs/2209.07753
- PerAct: Perceiver-Actor, a Multi-Task Transformer for Robotic Manipulation — University of Washington / NVIDIA · 2022-09 · paper · [manipulation] — https://arxiv.org/abs/2209.05451
- Deep Whole-Body Control: Learning a Unified Policy for Manipulation and Locomotion — Carnegie Mellon University · 2022-10 · paper · [locomotion-nav] — https://arxiv.org/abs/2210.10044
- GNM: A General Navigation Model to Drive Any Robot — UC Berkeley, Toyota Motor North America · 2022-10 · paper · [locomotion-nav] — https://arxiv.org/abs/2210.03370
- Interactive Language: Talking to Robots in Real Time — Google (Robotics at Google) · 2022-10 · paper · [manipulation] — https://arxiv.org/abs/2210.06407
- Real-World Robot Learning with Masked Visual Pre-training — UC Berkeley · 2022-10 · paper · [data] — https://arxiv.org/abs/2210.03109
- VIMA: General Robot Manipulation with Multimodal Prompts — Stanford / NVIDIA / Macalester / Caltech / Tsinghua / UT Austin · 2022-10 · paper · [benchmark] — https://arxiv.org/abs/2210.03094
- RT-1: Robotics Transformer for Real-World Control at Scale — Google (Robotics at Google / Everyday Robots) · 2022-12 · paper · [data] — https://arxiv.org/abs/2212.06817
- Walk These Ways: Tuning Robot Control for Generalization with Multiplicity of Behavior — MIT (Improbable AI) · 2022-12 · paper · [locomotion-nav] — https://arxiv.org/abs/2212.03238
2023(48 条)
- DreamWaQ: Learning Robust Quadrupedal Locomotion with Implicit Terrain Imagination via Deep Reinforcement Learning — KAIST · 2023-01 · paper · [locomotion-nav] — https://arxiv.org/abs/2301.10602
- Orbit: A Unified Simulation Framework for Interactive Robot Learning Environments — NVIDIA / ETH Zurich · 2023-01 · paper · [sim-infra] — https://arxiv.org/abs/2301.04195
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills — UC San Diego (Hao Su Lab) / Tsinghua University · 2023-02 · paper · [sim-infra] — https://arxiv.org/abs/2302.04659
- Scaling Robot Learning with Semantically Imagined Experience — Google (Robotics at Google / Google Research) · 2023-02 · paper · [data] — https://arxiv.org/abs/2302.11550
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion — Columbia University / Toyota Research Institute · 2023-03 · paper · [manipulation] — https://arxiv.org/abs/2303.04137
- PaLM-E: An Embodied Multimodal Language Model — Google (Robotics at Google / Google Research) / TU Berlin · 2023-03 · paper · [vla] — https://arxiv.org/abs/2303.03378
- Real-World Humanoid Locomotion with Reinforcement Learning — UC Berkeley · 2023-03 · paper · [humanoid] — https://arxiv.org/abs/2303.03381
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence? — Meta AI (FAIR) · 2023-03 · paper · [data] — https://arxiv.org/abs/2303.18240
- ACT) — Stanford University · 2023-04 · paper · [data] — https://arxiv.org/abs/2304.13705
- Neural Volumetric Memory for Visual Locomotion Control — UC San Diego · 2023-04 · paper · [locomotion-nav] — https://arxiv.org/abs/2304.01201
- Barkour: Benchmarking Animal-level Agility with Quadruped Robots — Google DeepMind · 2023-05 · paper · [benchmark] — https://arxiv.org/abs/2305.14654
- Perpetual Humanoid Control for Real-time Simulated Avatars — Meta Reality Labs Research / CMU · 2023-05 · paper · [humanoid] — https://arxiv.org/abs/2305.06456
- ANYmal Parkour: Learning Agile Navigation for Quadrupedal Robots — ETH Zurich · 2023-06 · paper · [locomotion-nav] — https://arxiv.org/abs/2306.14874
- Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation — Carnegie Mellon University · 2023-06 · paper · [manipulation] — https://arxiv.org/abs/2306.17817
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning — UT Austin / Sony AI / Tsinghua · 2023-06 · paper · [benchmark] — https://arxiv.org/abs/2306.03310
- RVT: Robotic View Transformer for 3D Object Manipulation — NVIDIA · 2023-06 · paper · [manipulation] — https://arxiv.org/abs/2306.14896
- RoboCat: A Self-Improving Foundation Agent for Robotic Manipulation — Google DeepMind · 2023-06 · paper · [manipulation] — https://arxiv.org/abs/2306.11706
- ViNT: A Foundation Model for Visual Navigation — UC Berkeley · 2023-06 · paper · [locomotion-nav] — https://arxiv.org/abs/2306.14846
- An Extensible, Data-Oriented Architecture for High-Performance, Many-World Simulation — Stanford University / Georgia Institute of Technology · 2023-07 · paper · [sim-infra] — https://madrona-engine.github.io/shacklett_siggraph23.pdf
- AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System — NVIDIA + UC San Diego · 2023-07 · paper · [data] — https://arxiv.org/abs/2307.04577
- RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot — Shanghai Jiao Tong University · 2023-07 · paper · [data] — https://arxiv.org/abs/2307.00595
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control — Google DeepMind · 2023-07 · paper · [vla] — https://arxiv.org/abs/2307.15818
- Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition — Columbia University + Google DeepMind · 2023-07 · paper · [manipulation] — https://arxiv.org/abs/2307.14535
- BridgeData V2: A Dataset for Robot Learning at Scale — UC Berkeley · 2023-08 · paper · [data] — https://arxiv.org/abs/2308.12952
- DTC: Deep Tracking Control — ETH Zurich · 2023-09 · paper · [locomotion-nav] — https://arxiv.org/abs/2309.15462
- Extreme Parkour with Legged Robots — CMU · 2023-09 · paper · [locomotion-nav] — https://arxiv.org/abs/2309.14341
- GELLO: A General, Low-Cost, and Intuitive Teleoperation Framework for Robot Manipulators — UC Berkeley · 2023-09 · paper · [data] — https://arxiv.org/abs/2309.13037
- Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions — Google DeepMind · 2023-09 · paper · [manipulation] — https://arxiv.org/abs/2309.10150
- RoboSet) — Carnegie Mellon University + Meta AI (FAIR) · 2023-09 · paper · [data] — https://arxiv.org/abs/2309.01918
- Robot Parkour Learning — Shanghai Qi Zhi Institute, Stanford, ShanghaiTech, CMU, Tsinghua · 2023-09 · paper · [locomotion-nav] — https://arxiv.org/abs/2309.05665
- Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots — Meta FAIR · 2023-10 · paper · [sim-infra] — https://arxiv.org/abs/2310.13724
- MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations — NVIDIA + UT Austin · 2023-10 · paper · [data] — https://arxiv.org/abs/2310.17596
- NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration — UC Berkeley · 2023-10 · paper · [locomotion-nav] — https://arxiv.org/abs/2310.07896
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models — Google DeepMind + 33 academic labs · 2023-10 · paper · [data] — https://arxiv.org/abs/2310.08864
- ViPlanner: Visual Semantic Imperative Learning for Local Navigation — ETH Zurich, NVIDIA · 2023-10 · paper · [locomotion-nav] — https://arxiv.org/abs/2310.00982
- cuRobo: Parallelized Collision-Free Minimum-Jerk Robot Motion Generation — NVIDIA · 2023-10 · tech-report · [sim-infra] — https://arxiv.org/abs/2310.17274
- ChainedDiffuser: Unifying Trajectory Diffusion and Keypose Prediction for Robotic Manipulation — Carnegie Mellon University · 2023-11 · paper · [manipulation] — https://proceedings.mlr.press/v229/xian23a.html
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives — Meta FAIR + Project Aria + 15-university consortium · 2023-11 · paper · [data] — https://arxiv.org/abs/2311.18259
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches — Google DeepMind · 2023-11 · paper · [manipulation] — https://arxiv.org/abs/2311.01977
- RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation — CMU, Tsinghua IIIS, MIT CSAIL, UMass Amherst, MIT-IBM AI Lab · 2023-11 · paper · [sim-infra] — https://arxiv.org/abs/2311.01455
- RoboVQA: Multimodal Long-Horizon Reasoning for Robotics — Google DeepMind · 2023-11 · paper · [benchmark] — https://arxiv.org/abs/2311.00899
- Robot Learning in the Era of Foundation Models: A Survey — Tongji University · Shaanxi University of Science and Technology · 2023-11 · paper · [survey] — https://arxiv.org/abs/2311.14379
- Any-point Trajectory Modeling for Policy Learning — UC Berkeley · 2023-12 · paper · [manipulation] — https://arxiv.org/abs/2401.00025
- Foundation Models in Robotics: Applications, Challenges, and the Future — Stanford / Princeton / UT Austin / NVIDIA / Google DeepMind / TU Berlin / SJTU / Scaled Foundations · 2023-12 · paper · [survey] — https://arxiv.org/abs/2312.07843
- MuJoCo XLA (MJX) — DeepMind · 2023-12 · tech-report · [sim-infra] — https://mujoco.readthedocs.io/en/stable/mjx.html
- SARA-RT: Scaling up Robotics Transformers with Self-Adaptive Robust Attention — Google DeepMind · 2023-12 · paper · [vla] — https://arxiv.org/abs/2312.01990
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation — ByteDance Research · 2023-12 · paper · [manipulation] — https://arxiv.org/abs/2312.13139
- VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation — Boston Dynamics AI Institute, Georgia Tech · 2023-12 · paper · [locomotion-nav] — https://arxiv.org/abs/2312.03275
2024(98 条)
- Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion — Carnegie Mellon University; ETH Zürich · 2024-01 · paper · [locomotion-nav] — https://arxiv.org/abs/2401.17583
- AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents — Google DeepMind · 2024-01 · paper · [data] — https://arxiv.org/abs/2401.12963
- General Flow as Foundation Affordance for Scalable Robot Learning — Tsinghua University · 2024-01 · paper · [manipulation] — https://arxiv.org/abs/2401.11439
- Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation — Stanford University · 2024-01 · paper · [data] — https://arxiv.org/abs/2401.02117
- 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations — Carnegie Mellon University · 2024-02 · paper · [manipulation] — https://arxiv.org/abs/2402.10885
- AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies — University of Texas at Austin · 2024-02 · paper · [manipulation] — https://arxiv.org/abs/2402.04292
- Expressive Whole-Body Control for Humanoid Robots — UC San Diego · 2024-02 · paper · [humanoid] — https://arxiv.org/abs/2402.16796
- NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation — Peking University (PKU-EPIC), BAAI, Galbot · 2024-02 · paper · [locomotion-nav] — https://arxiv.org/abs/2402.15852
- PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs — Google DeepMind · 2024-02 · paper · [vla] — https://arxiv.org/abs/2402.07872
- THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation — University of Washington, NVIDIA · 2024-02 · paper · [benchmark] — https://arxiv.org/abs/2402.08191
- Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots — Stanford + Columbia + Toyota Research Institute · 2024-02 · paper · [data] — https://arxiv.org/abs/2402.10329
- 3D Diffusion Policy (DP3): Generalizable Visuomotor Policy Learning via Simple 3D Representations — Shanghai Qi Zhi Institute / Tsinghua University / Shanghai Jiao Tong University · 2024-03 · paper · [manipulation] — https://arxiv.org/abs/2403.03954
- BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation — Stanford (Vision & Learning Lab) · 2024-03 · paper · [benchmark] — https://arxiv.org/abs/2403.09227
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset — Stanford / UC Berkeley / TRI + 13-lab consortium · 2024-03 · paper · [data] — https://arxiv.org/abs/2403.12945
- DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation — Stanford University · 2024-03 · paper · [data] — https://arxiv.org/abs/2403.07788
- Hierarchical Diffusion Policy for Kinematics-Aware Multi-Task Robotic Manipulation — Dyson Robot Learning Lab · 2024-03 · paper · [manipulation] — https://arxiv.org/abs/2403.03890
- HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation — UC Berkeley · 2024-03 · paper · [benchmark] — https://arxiv.org/abs/2403.10506
- Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation (H2O) — Carnegie Mellon University (LeCAR Lab) · 2024-03 · paper · [humanoid] — https://arxiv.org/abs/2403.04436
- RT-H: Action Hierarchies Using Language — Google DeepMind · 2024-03 · paper · [vla] — https://arxiv.org/abs/2403.01823
- RT-Sketch: Goal-Conditioned Imitation Learning from Hand-Drawn Sketches — Stanford University / Google DeepMind / Google Intrinsic · 2024-03 · paper · [manipulation] — https://arxiv.org/abs/2403.02709
- VQ-BeT: Behavior Generation with Latent Actions — New York University · 2024-03 · paper · [manipulation] — https://arxiv.org/abs/2403.03181
- Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers — Google DeepMind · 2024-03 · paper · [manipulation] — https://arxiv.org/abs/2403.12943
- RISE: 3D Perception Makes Real-World Robot Imitation Simple and Effective — Shanghai Jiao Tong University · 2024-04 · paper · [manipulation] — https://arxiv.org/abs/2404.12281
- Towards Generalist Robot Learning from Internet Video: A Survey — University College London · Weco AI · MIT · 2024-04 · paper · [survey] — https://arxiv.org/abs/2404.19664
- A Survey on Vision-Language-Action Models for Embodied AI — CUHK · Huawei Noah’s Ark Lab · 2024-05 · paper · [survey] — https://arxiv.org/abs/2405.14093
- Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation — Stanford University / Princeton University · 2024-05 · paper · [manipulation] — https://arxiv.org/abs/2405.07503
- Evaluating Real-World Robot Manipulation Policies in Simulation (SIMPLER) — UC San Diego, Stanford, UC Berkeley, Google DeepMind · 2024-05 · paper · [benchmark] — https://arxiv.org/abs/2405.05941
- Octo: An Open-Source Generalist Robot Policy — UC Berkeley, Stanford, CMU, Google DeepMind · 2024-05 · paper · [vla] — https://arxiv.org/abs/2405.12213
- Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning — Dyson Robot Learning Lab / Imperial College London · 2024-05 · paper · [manipulation] — https://arxiv.org/abs/2405.18196
- SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation — Tsinghua University / Shanghai AI Laboratory · 2024-05 · paper · [manipulation] — https://arxiv.org/abs/2405.19586
- Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation — Carnegie Mellon University, University of Washington, Meta · 2024-05 · paper · [data] — https://arxiv.org/abs/2405.01527
- 3D-MVP: 3D Multiview Pretraining for Robotic Manipulation — NVIDIA · 2024-06 · paper · [manipulation] — https://arxiv.org/abs/2406.18158
- BAKU: An Efficient Transformer for Multi-Task Policy Learning — New York University · 2024-06 · paper · [manipulation] — https://arxiv.org/abs/2406.07539
- CityNav: Language-Goal Aerial Navigation Dataset with Geographic Information — Institute of Science Tokyo / University of Tokyo / NII / ATR / RIKEN AIP / Kyoto University / Sony Semiconductor Solutions · 2024-06 · paper · [locomotion-nav] — https://arxiv.org/abs/2406.14240
- HumanPlus: Humanoid Shadowing and Imitation from Humans — Stanford University · 2024-06 · paper · [data] — https://arxiv.org/abs/2406.10454
- Humanoid Parkour Learning — Shanghai Qi Zhi Institute / ShanghaiTech University / Tsinghua University · 2024-06 · paper · [humanoid] — https://arxiv.org/abs/2406.10759
- ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation — Tsinghua University (SIGS) / Shanghai AI Laboratory / Carnegie Mellon University · 2024-06 · paper · [manipulation] — https://arxiv.org/abs/2406.01586
- ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data — Stanford University + Columbia University + Toyota Research Institute · 2024-06 · paper · [manipulation] — https://arxiv.org/abs/2406.19464
- OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning — Carnegie Mellon University (LeCAR Lab) · 2024-06 · paper · [humanoid] — https://arxiv.org/abs/2406.08858
- OpenVLA: An Open-Source Vision-Language-Action Model — Stanford, UC Berkeley, Toyota Research Institute, Google DeepMind · 2024-06 · paper · [vla] — https://arxiv.org/abs/2406.09246
- RVT-2: Learning Precise Manipulation from Few Demonstrations — NVIDIA · 2024-06 · paper · [manipulation] — https://arxiv.org/abs/2406.08545
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots — UT Austin, NVIDIA · 2024-06 · paper · [sim-infra] — https://arxiv.org/abs/2406.02523
- SGRv2: Leveraging Locality to Boost Sample Efficiency in Robotic Manipulation — Tsinghua University / Shanghai Qi Zhi Institute · 2024-06 · paper · [manipulation] — https://arxiv.org/abs/2406.10615
- Streaming Diffusion Policy: Fast Policy Synthesis with Variable Noise Diffusion Models — NTNU / Harvard University · 2024-06 · paper · [manipulation] — https://arxiv.org/abs/2406.04806
- WoCoCo: Learning Whole-Body Humanoid Control with Sequential Contacts — Carnegie Mellon University (LeCAR Lab) · 2024-06 · paper · [humanoid] — https://arxiv.org/abs/2406.06005
- Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI — SYSU HCP Lab · Peng Cheng Laboratory · PKU · 2024-07 · paper · [survey] — https://arxiv.org/abs/2407.06886
- BiGym: A Demo-Driven Mobile Bi-Manual Manipulation Benchmark — Dyson Robot Learning Lab, UCL · 2024-07 · paper · [benchmark] — https://arxiv.org/abs/2407.07788
- Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning — UC San Diego + University of Hong Kong · 2024-07 · paper · [data] — https://arxiv.org/abs/2407.03162
- EquiBot: SIM(3)-Equivariant Diffusion Policy for Generalizable and Data Efficient Learning — Stanford University · 2024-07 · paper · [manipulation] — https://arxiv.org/abs/2407.01479
- Equivariant Diffusion Policy — Northeastern University / Boston Dynamics AI Institute · 2024-07 · paper · [manipulation] — https://arxiv.org/abs/2407.01812
- Flow as the Cross-Domain Manipulation Interface — Stanford University, Columbia University, JP Morgan AI Research, Carnegie Mellon University · 2024-07 · paper · [manipulation] — https://arxiv.org/abs/2407.15208
- GRUtopia: Dream General Robots in a City at Scale — Shanghai AI Laboratory (OpenRobotLab) · 2024-07 · paper · [sim-infra] — https://arxiv.org/abs/2407.10943
- Generative Image as Action Models — Dyson Robot Learning Lab · 2024-07 · paper · [manipulation] — https://arxiv.org/abs/2407.07875
- MetaUrban: An Embodied AI Simulation Platform for Urban Micromobility — UCLA (VAIL Lab, Bolei Zhou) / MetaDriverse · 2024-07 · paper · [sim-infra] — https://arxiv.org/abs/2407.08725
- Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs — Google DeepMind · 2024-07 · paper · [locomotion-nav] — https://arxiv.org/abs/2407.07775
- Open-TeleVision: Teleoperation with Immersive Active Visual Feedback — UC San Diego + MIT · 2024-07 · paper · [data] — https://arxiv.org/abs/2407.01512
- Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning — UC Berkeley / Carnegie Mellon University · 2024-07 · paper · [manipulation] — https://arxiv.org/abs/2407.01531
- UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers — Stanford + Columbia + Google DeepMind · 2024-07 · paper · [data] — https://arxiv.org/abs/2407.10353
- All Robots in One: A New Standard and Unified Dataset for Versatile, General-Purpose Embodied Agents — Pengcheng Laboratory · Southern University of Science and Technology · Sun Yat-sen University · 2024-08 · paper · [data] — https://arxiv.org/abs/2408.10899
- Learning Multi-Modal Whole-Body Control for Real-World Humanoid Robots — Oregon State University · 2024-08 · paper · [humanoid] — https://arxiv.org/abs/2408.07295
- Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation — UC Berkeley, Carnegie Mellon University · 2024-08 · paper · [locomotion-nav] — https://arxiv.org/abs/2408.11812
- DPPO: Diffusion Policy Policy Optimization — Princeton University · 2024-09 · paper · [manipulation] — https://arxiv.org/abs/2409.00588
- DemoStart: Demonstration-led Auto-curriculum Applied to Sim-to-real with Multi-fingered Robots — Google DeepMind · 2024-09 · paper · [manipulation] — https://arxiv.org/abs/2409.06613
- HPT: Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers — MIT · 2024-09 · paper · [manipulation] — https://arxiv.org/abs/2409.20537
- Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models — Stanford University, UC San Diego · 2024-09 · paper · [locomotion-nav] — https://arxiv.org/abs/2410.00231
- ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation — Stanford University · 2024-09 · paper · [manipulation] — https://arxiv.org/abs/2409.01652
- TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation — East China Normal University, Midea Group AI Lab, Shanghai University, Syracuse University, Beijing Innovation Center of Humanoid Robotics · 2024-09 · paper · [vla] — https://arxiv.org/abs/2409.12514
- ALOHA Unleashed: A Simple Recipe for Robot Dexterity — Google DeepMind · 2024-10 · paper · [manipulation] — https://arxiv.org/abs/2410.13126
- Data Scaling Laws in Imitation Learning for Robotic Manipulation — Tsinghua University + Shanghai Qi Zhi Institute + Shanghai AI Lab · 2024-10 · paper · [data] — https://arxiv.org/abs/2410.18647
- DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning — NVIDIA + UT Austin + UC San Diego · 2024-10 · paper · [data] — https://arxiv.org/abs/2410.24185
- EgoMimic: Scaling Imitation Learning via Egocentric Video — Georgia Institute of Technology · 2024-10 · paper · [data] — https://arxiv.org/abs/2410.24221
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation — ByteDance Research · 2024-10 · tech-report · [vla] — https://gr2-manipulation.github.io/
- HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots — NVIDIA · 2024-10 · paper · [humanoid] — https://arxiv.org/abs/2410.21229
- Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions — University of Texas at Austin (RPL); NVIDIA Research · 2024-10 · paper · [humanoid] — https://arxiv.org/abs/2410.12773
- Latent Action Pretraining from Videos (LAPA) — KAIST · University of Washington · Microsoft Research · NVIDIA · AI2 · 2024-10 · paper · [vla] — https://arxiv.org/abs/2410.11758
- LeLaN: Learning A Language-Conditioned Navigation Policy from In-the-Wild Videos — UC Berkeley, Toyota Motor North America · 2024-10 · paper · [locomotion-nav] — https://arxiv.org/abs/2410.03603
- ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI — UC San Diego (Hao Su Lab) · 2024-10 · paper · [sim-infra] — https://arxiv.org/abs/2410.00425
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation — Tsinghua (TSAIL) · 2024-10 · paper · [vla] — https://arxiv.org/abs/2410.07864
- Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy — Inria Paris (WILLOW), ENS/PSL · 2024-10 · paper · [benchmark] — https://arxiv.org/abs/2410.01345
- iDP3: Generalizable Humanoid Manipulation with 3D Diffusion Policies — Stanford University · 2024-10 · paper · [humanoid] — https://arxiv.org/abs/2410.10803
- π0: A Vision-Language-Action Flow Model for General Robot Control — Physical Intelligence · 2024-10 · paper · [vla] — https://arxiv.org/abs/2410.24164
- CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos — New York University · 2024-11 · paper · [locomotion-nav] — https://arxiv.org/abs/2411.17820
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation — Microsoft Research Asia, Tsinghua University, USTC, Institute of Microelectronics (CAS) · 2024-11 · paper · [vla] — https://arxiv.org/abs/2411.19650
- DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution — Tsinghua University · ByteDance Research · 2024-11 · paper · [vla] — https://arxiv.org/abs/2411.02359
- WildLMa: Long Horizon Loco-Manipulation in the Wild — UC San Diego · 2024-11 · paper · [locomotion-nav] — https://arxiv.org/abs/2411.15131
- Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning — Midea Group / East China Normal University · 2024-12 · paper · [vla] — https://arxiv.org/abs/2412.03293
- Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning — Karlsruhe Institute of Technology · 2024-12 · paper · [manipulation] — https://arxiv.org/abs/2412.12953
- ExBody2: Advanced Expressive Humanoid Whole-Body Control — UC San Diego / UC Berkeley / MIT / NVIDIA · 2024-12 · paper · [humanoid] — https://arxiv.org/abs/2412.13196
- Genesis: A Generative and Universal Physics Engine for Robotics and Embodied AI — Genesis Authors(CMU / Stanford 等 20+ 机构学术协作,后成立 Genesis AI) · 2024-12 · model-card · [sim-infra] — https://github.com/Genesis-Embodied-AI/genesis-world
- Learning from Massive Human Videos for Universal Humanoid Pose Control — USC / UC Berkeley / Toyota Research Institute · 2024-12 · paper · [humanoid] — https://arxiv.org/abs/2412.14172
- Mobile-TeleVision: Predictive Motion Priors for Humanoid Whole-Body Control — UC San Diego / MIT · 2024-12 · paper · [humanoid] — https://arxiv.org/abs/2412.07773
- Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos — University of Hong Kong, Tencent ARC Lab, Chinese University of Hong Kong, UC Berkeley · 2024-12 · paper · [manipulation] — https://arxiv.org/abs/2412.04445
- NaVILA: Legged Robot Vision-Language-Action Model for Navigation — UC San Diego / USC / NVIDIA · 2024-12 · paper · [locomotion-nav] — https://arxiv.org/abs/2412.04453
- Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation — Shanghai AI Laboratory, Peking University, Chinese University of Hong Kong · 2024-12 · paper · [manipulation] — https://arxiv.org/abs/2412.15109
- RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation — Beijing Innovation Center of Humanoid Robotics · Peking University · BAAI · 2024-12 · paper · [data] — https://arxiv.org/abs/2412.13877
- Towards Generalist Robot Policies: What Matters in Building Vision-Language-Action Models — ByteDance Research, Tsinghua University, CASIA MAIS-NLPR, Shanghai Jiao Tong University, National University of Singapore · 2024-12 · paper · [vla] — https://arxiv.org/abs/2412.14058
- Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks — Peking University (PKU-EPIC), Galbot, Beijing Academy of Artificial Intelligence · 2024-12 · paper · [locomotion-nav] — https://arxiv.org/abs/2412.06224
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks — Fudan University · 2024-12 · paper · [benchmark] — https://arxiv.org/abs/2412.18194
2025(93 条)
- FAST: Efficient Action Tokenization for Vision-Language-Action Models — Physical Intelligence · 2025-01 · paper · [vla] — https://arxiv.org/abs/2501.09747
- SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model — Shanghai AI Laboratory / ShanghaiTech · 2025-01 · paper · [vla] — https://arxiv.org/abs/2501.15830
- Universal Actions for Enhanced Embodied Foundation Models — AIR Tsinghua University / SenseTime Research / Peking University / Beijing University of Posts and Telecommunications / Shanghai AI Lab · 2025-01 · paper · [vla] — https://arxiv.org/abs/2501.10105
- ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills — Carnegie Mellon University (LeCAR Lab) / NVIDIA · 2025-02 · paper · [humanoid] — https://arxiv.org/abs/2502.01143
- ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model — Midea Group / East China Normal University / Shanghai University / Beijing Innovation Center of Humanoid Robotics / Tsinghua University · 2025-02 · paper · [vla] — https://arxiv.org/abs/2502.14420
- CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World — Peking University · 2025-02 · paper · [manipulation] — https://arxiv.org/abs/2502.08449
- DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control — Midea Group / East China Normal University · 2025-02 · paper · [vla] — https://arxiv.org/abs/2502.05855
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents — University of Illinois Urbana-Champaign, Northwestern University, University of Toronto, TTIC · 2025-02 · paper · [benchmark] — https://arxiv.org/abs/2502.09560
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT) — Stanford University · 2025-02 · paper · [vla] — https://arxiv.org/abs/2502.19645
- HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit — Shanghai AI Lab (OpenRobotLab) · 2025-02 · paper · [humanoid] — https://arxiv.org/abs/2502.13013
- Helix: A Vision-Language-Action Model for Generalist Humanoid Control — Figure AI · 2025-02 · blog · [vla] — https://www.figure.ai/news/helix
- HugWBC: A Unified and General Humanoid Whole-Body Controller for Versatile Locomotion — Shanghai Jiao Tong University / Shanghai AI Lab · 2025-02 · paper · [humanoid] — https://arxiv.org/abs/2502.03206
- Learning Getting-Up Policies for Real-World Humanoid Robots — University of Illinois Urbana-Champaign / Simon Fraser University · 2025-02 · paper · [humanoid] — https://arxiv.org/abs/2502.12152
- Magma: A Foundation Model for Multimodal AI Agents — Microsoft Research · 2025-02 · paper · [vla] — https://arxiv.org/abs/2502.13130
- MuJoCo Playground: An Open-Source Framework for GPU-Accelerated Robot Learning and Sim-to-Real Transfer — UC Berkeley / Google DeepMind / U Toronto / U Cambridge · 2025-02 · paper · [sim-infra] — https://arxiv.org/abs/2502.08844
- RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete — BAAI · 2025-02 · paper · [vla] — https://arxiv.org/abs/2502.21257
- Survey on Vision-Language-Action Models — Tactile Robotics Laboratory (Kazakhstan) · 2025-02 · paper · [survey] — https://arxiv.org/abs/2502.06851
- AgiBot Genie Sim: Open-Source Simulation Platform for Embodied Intelligence — AgiBot (智元机器人) · 2025-03 · tech-report · [sim-infra] — https://github.com/AgibotTech/genie_sim
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems — AgiBot (智元) · Shanghai AI Lab · HKU · Shanghai Innovation Institute · 2025-03 · paper · [data] — https://arxiv.org/abs/2503.06669
- Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills — Peking University / BAAI / BeingBeyond · 2025-03 · paper · [humanoid] — https://arxiv.org/abs/2503.12533
- Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy — Shanghai AI Lab / Zhejiang University / CUHK MMLab / Peking University / SenseTime Research / Tsinghua University / HKISI-CAS · 2025-03 · paper · [vla] — https://arxiv.org/abs/2503.19757
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots — NVIDIA · 2025-03 · tech-report · [vla] — https://arxiv.org/abs/2503.14734
- Gemini Robotics: Bringing AI into the Physical World — Google DeepMind · 2025-03 · tech-report · [vla] — https://arxiv.org/abs/2503.20020
- Humanoid Policy ~ Human Policy — UC San Diego + CMU + University of Washington + MIT · 2025-03 · paper · [data] — https://arxiv.org/abs/2503.13441
- Newton: An Open-Source GPU-Accelerated Physics Engine for Robotics — NVIDIA / Google DeepMind / Disney Research (Linux Foundation) · 2025-03 · blog · [sim-infra] — https://developer.nvidia.com/blog/announcing-newton-an-open-source-physics-engine-for-robotics-simulation/
- Dexterous Manipulation through Imitation Learning: A Survey — Tianjin University · Shandong University · CAS Institute of Automation · 2025-04 · paper · [survey] — https://arxiv.org/abs/2504.03515
- RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins — HKU, AgileX Robotics, Shanghai AI Laboratory, SZU, CASIA, UNC-Chapel Hill, GDIIST, SJTU · 2025-04 · paper · [benchmark] — https://arxiv.org/abs/2504.13059
- RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning — UC Berkeley + Peking University + USC + consortium (9 institutions) · 2025-04 · paper · [sim-infra] — https://arxiv.org/abs/2504.18904
- ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface — Tsinghua University + University of California, Berkeley · 2025-04 · paper · [data] — https://arxiv.org/abs/2504.06156
- π0.5: a VLA Model with Open-World Generalization — Physical Intelligence · 2025-04 · paper · [vla] — https://www.physicalintelligence.company/blog/pi05
- AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control — UC San Diego · 2025-05 · paper · [humanoid] — https://arxiv.org/abs/2505.03738
- ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge — Midea Group / East China Normal University · 2025-05 · paper · [vla] — https://arxiv.org/abs/2505.21906
- DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation — Columbia + Stanford · 2025-05 · paper · [data] — https://arxiv.org/abs/2505.21864
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models — NVIDIA · 2025-05 · paper · [data] — https://arxiv.org/abs/2505.12705
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video — Apple · 2025-05 · paper · [data] — https://arxiv.org/abs/2505.11709
- FALCON: Learning Force-Adaptive Humanoid Loco-Manipulation — Carnegie Mellon University (LeCAR Lab) · 2025-05 · paper · [humanoid] — https://arxiv.org/abs/2505.06776
- GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data — Galbot / Peking University · 2025-05 · paper · [manipulation] — https://arxiv.org/abs/2505.03233
- HuB: Learning Extreme Humanoid Balance — Tsinghua University / Shanghai Qi Zhi Institute / Shanghai AI Lab · 2025-05 · paper · [humanoid] — https://arxiv.org/abs/2505.07294
- Hume: Introducing System-2 Thinking in Visual-Language-Action Model — Shanghai AI Laboratory / SJTU / Zhejiang University / AgiBot · 2025-05 · paper · [vla] — https://arxiv.org/abs/2505.21432
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better — Physical Intelligence · 2025-05 · paper · [vla] — https://www.physicalintelligence.company/research/knowledge_insulation
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation — Westlake University / Zhejiang University · 2025-05 · paper · [vla] — https://arxiv.org/abs/2505.03912
- ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning — Tsinghua University · 2025-05 · paper · [manipulation] — https://arxiv.org/abs/2505.22094
- TWIST: Teleoperated Whole-Body Imitation System — Stanford University / Simon Fraser University · 2025-05 · paper · [humanoid] — https://arxiv.org/abs/2505.02833
- Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction — TeleAI (China Telecom) · 2025-05 · paper · [manipulation] — https://arxiv.org/abs/2505.24156
- TrackVLA: Embodied Visual Tracking in the Wild — Peking University (PKU-EPIC) / Galbot / Beihang University / Beijing Normal University / BAAI · 2025-05 · paper · [locomotion-nav] — https://arxiv.org/abs/2505.23189
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions — OpenDriveLab / AgiBot / The University of Hong Kong · 2025-05 · paper · [vla] — https://arxiv.org/abs/2505.06111
- A Survey of Behavior Foundation Model: Next-Generation Whole-Body Control System of Humanoid Robots — HK PolyU · LimX Dynamics · Eastern Institute of Technology (Ningbo) · HKU · EPFL · 2025-06 · paper · [survey] — https://arxiv.org/abs/2506.20487
- CLONE: Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks — Beijing Institute of Technology / BIGAI / Peking University · 2025-06 · paper · [humanoid] — https://arxiv.org/abs/2506.08931
- DYNA-1: A Commercial-Grade Autonomous Dexterous Model — Dyna Robotics · 2025-06 · blog · [manipulation] — https://www.dyna.co/research/dyna-1
- GMT: General Motion Tracking for Humanoid Whole-Body Control — UC San Diego / Simon Fraser University · 2025-06 · paper · [humanoid] — https://arxiv.org/abs/2506.14770
- GR00T N1.5: An Improved Open Foundation Model for Generalist Humanoid Robots — NVIDIA · 2025-06 · blog · [vla] — https://research.nvidia.com/labs/gear/gr00t-n1_5/
- GenManip: LLM-driven Simulation for Generalizable Instruction-Following Manipulation — Shanghai AI Laboratory, Xi’an Jiaotong University, Zhejiang University, Nanjing University · 2025-06 · paper · [benchmark] — https://arxiv.org/abs/2506.10966
- KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills — Institute of Artificial Intelligence (TeleAI), China Telecom / Shanghai Jiao Tong University · 2025-06 · paper · [humanoid] — https://arxiv.org/abs/2506.12851
- LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction — UC Berkeley / Carnegie Mellon University · 2025-06 · paper · [humanoid] — https://arxiv.org/abs/2506.13751
- RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies — UC Berkeley / Stanford / UW / UT Austin (DROID network) · 2025-06 · paper · [benchmark] — https://arxiv.org/abs/2506.18123
- RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation — Beihang University, National University of Singapore, Shanghai Jiao Tong University · 2025-06 · paper · [benchmark] — https://arxiv.org/abs/2506.06677
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation — SJTU · HKU MMLab · Shanghai AI Lab · SZU · THU · TeleAI · FDU · USTC · SUSTech · SYSU · CSU · NEU · NJU · Lumina EAI (硬件支持 AgileX Robotics,软件支持 D-Robotics,AIGC 支持 Deemos) · 2025-06 · paper · [benchmark] — https://arxiv.org/abs/2506.18088
- SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending — University of Southern California / Stanford University / UC Berkeley / Peking University · 2025-06 · paper · [humanoid] — https://arxiv.org/abs/2506.09366
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics — Hugging Face · 2025-06 · paper · [vla] — https://arxiv.org/abs/2506.01844
- WorldVLA: Towards Autoregressive Action World Model — Alibaba DAMO Academy · 2025-06 · paper · [vla] — https://arxiv.org/abs/2506.21539
- A Survey: Learning Embodied Intelligence from Physical Simulators and World Models — Nanjing University 领衔的多机构联合团队(NJU3DV-LoongGroup) · 2025-07 · paper · [survey] — https://arxiv.org/abs/2507.00917
- GR-3 Technical Report — ByteDance Seed · 2025-07 · tech-report · [vla] — https://seed.bytedance.com/GR3
- RoboBrain 2.0 Technical Report — BAAI · 2025-07 · tech-report · [vla] — https://arxiv.org/abs/2507.02029
- Skild Brain: An Omni-Bodied Robotic Foundation Model — Skild AI · 2025-07 · blog · [vla] — https://www.skild.ai/blogs/building-the-general-purpose-robotic-brain
- The Developments and Challenges towards Dexterous and Embodied Robotic Manipulation: A Survey — Zhejiang University · 2025-07 · paper · [survey] — https://arxiv.org/abs/2507.11840
- UniTracker: Learning Universal Whole-Body Motion Tracker for Humanoid Robots — Shanghai Jiao Tong University / Shanghai AI Lab · 2025-07 · paper · [humanoid] — https://arxiv.org/abs/2507.07356
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models — Microsoft Research · 2025-07 · paper · [vla] — https://arxiv.org/abs/2507.23682
- BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion — UC Berkeley (Hybrid Robotics) × Stanford · 2025-08 · paper · [humanoid] — https://beyondmimic.github.io/
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey — Harbin Institute of Technology (Shenzhen) · 2025-08 · paper · [survey] — https://arxiv.org/abs/2508.13073
- MolmoAct: Action Reasoning Models that can Reason in Space — Allen Institute for AI (Ai2) · 2025-08 · paper · [vla] — https://arxiv.org/abs/2508.07917
- Survey of Vision-Language-Action Models for Embodied Manipulation — Institute of Automation, Chinese Academy of Sciences · Beijing Zhongke Huiling Robot Technology · University of Chinese Academy of Sciences · 2025-08 · paper · [survey] — https://arxiv.org/abs/2508.15201
- Galaxea Open-World Dataset and G0 Dual-System VLA Model — Galaxea (星海图) · 2025-09 · paper · [vla] — https://arxiv.org/abs/2509.00576
- Gemini Robotics 1.5 and Gemini Robotics-ER 1.5 — Google DeepMind · 2025-09 · tech-report · [vla] — https://arxiv.org/abs/2510.03342
- MV-UMI: A Scalable Multi-View Interface for Cross-Embodiment Learning — New York University Abu Dhabi · 2025-09 · paper · [data] — https://arxiv.org/abs/2509.18757
- OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction — Amazon FAR × MIT × UC Berkeley × Stanford × CMU · 2025-09 · paper · [humanoid] — https://omniretarget.github.io/
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation — Alibaba DAMO Academy · 2025-09 · paper · [vla] — https://arxiv.org/abs/2509.15212
- VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation — Stanford University · 2025-09 · paper · [humanoid] — https://arxiv.org/abs/2509.20322
- WALL-OSS: Igniting VLMs toward the Embodied Space — X Square Robot (自变量机器人) · 2025-09 · paper · [vla] — https://arxiv.org/abs/2509.11766
- 1X NEO + Redwood AI (learned world-model policy) — 1X Technologies · 2025-10 · blog · [humanoid] — https://www.1x.tech/neo
- A Comprehensive Survey on World Models for Embodied AI — Nankai University / Tianjin University of Technology / UESTC · 2025-10 · paper · [survey] — https://arxiv.org/abs/2510.16732
- A Survey on Efficient Vision-Language-Action Models — Tongji University / Southwest Jiaotong University / UESTC / University of Trento · 2025-10 · paper · [survey] — https://arxiv.org/abs/2510.24795
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy — Shanghai AI Laboratory · 2025-10 · paper · [vla] — https://arxiv.org/abs/2510.13778
- PhysHSI: Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System — Shanghai AI Laboratory × HKUST · 2025-10 · paper · [humanoid] — https://why618188.github.io/physhsi/
- ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning — Amazon FAR (Frontier AI & Robotics) / USC / Stanford University / UC Berkeley / CMU · 2025-10 · paper · [humanoid] — https://arxiv.org/abs/2510.05070
- TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking — Peking University / Galbot / USTC / BAAI / Beihang University / SUSTech / Beijing Normal University · 2025-10 · paper · [locomotion-nav] — https://arxiv.org/abs/2510.07134
- ACE-F: A Cross Embodiment Foldable System with Force Feedback for Dexterous Teleoperation — UC San Diego · 2025-11 · paper · [data] — https://arxiv.org/abs/2511.20887
- BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning — Carnegie Mellon University / Meta (FAIR) · 2025-11 · paper · [humanoid] — https://arxiv.org/abs/2511.04131
- In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task Data — UC San Diego · 2025-11 · paper · [data] — https://arxiv.org/abs/2511.15704
- RynnVLA-002: A Unified Vision-Language-Action and World Model — Alibaba DAMO Academy · 2025-11 · paper · [vla] — https://arxiv.org/abs/2511.17502
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control — NVIDIA · 2025-11 · paper · [humanoid] — https://arxiv.org/abs/2511.07820
- TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System — Stanford University / Amazon FAR · 2025-11 · paper · [data] — https://arxiv.org/abs/2511.02832
- π*0.6: a VLA That Learns From Experience (RECAP) — Physical Intelligence · 2025-11 · tech-report · [vla] — https://www.pi.website/blog/pistar06
- WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation Control — AgiBot (智元) / OpenDriveLab & MMLab at HKU / Fudan University / Shanghai Innovation Institute · 2025-12 · paper · [vla] — https://arxiv.org/abs/2512.11047
2026(12 条)
- InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation — Shanghai AI Laboratory · 2026-01 · paper · [vla] — https://arxiv.org/abs/2601.02456
- RoboBrain 2.5: Depth in Sight, Time in Mind — BAAI · 2026-01 · tech-report · [vla] — https://arxiv.org/abs/2601.14352
- RoboReward: General-Purpose Vision-Language Reward Models for Robotics — Stanford University, UC Berkeley · 2026-01 · paper · [benchmark] — https://arxiv.org/abs/2601.00675
- One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies — University of Michigan / Lehigh University / Tsinghua University · 2026-03 · paper · [manipulation] — https://arxiv.org/abs/2603.12480
- SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation — Sichuan University / UESTC / Dexmal Inc. · 2026-03 · paper · [manipulation] — https://arxiv.org/abs/2603.05117
- Ψ₀: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation — USC Physical Superintelligence (PSI) Lab · 2026-03 · paper · [humanoid] — https://arxiv.org/abs/2603.12263
- R3D: Revisiting 3D Policy Learning — Zhejiang University / ShanghaiTech University · 2026-04 · paper · [manipulation] — https://arxiv.org/abs/2604.15281
- π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities — Physical Intelligence · 2026-04 · tech-report · [vla] — https://www.physicalintelligence.company/blog/pi07
- Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation — Nanjing University / Institute of Automation CAS / Beihang University / BMW (Nanjing) Information Technology / University of Rochester · 2026-05 · paper · [locomotion-nav] — https://arxiv.org/abs/2605.27582
- Wall-OSS-0.5 Technical Report: Pretrain Once, Act Anywhere — X Square Robot (自变量机器人) · 2026-05 · tech-report · [vla] — https://arxiv.org/abs/2605.30877
- Galaxea G0.5 Technical Report — Galaxea (星海图) · 2026-06 · tech-report · [vla] — https://opengalaxea.github.io/G05/
- InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization — Shanghai AI Laboratory · 2026-07 · paper · [vla] — https://arxiv.org/abs/2607.04988