AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: vision-language-action
此标签下有9条笔记。
2026年7月16日
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
rt-2
vla
vision-language-action
pali-x
palm-e
co-fine-tuning
action-as-text-token
emergent-reasoning
chain-of-thought
closed-loop-control
2026年7月16日
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
VLN
vision-language-action
legged-robot
quadruped
humanoid
locomotion-RL
VILA
LiDAR
isaac-sim
sim2real
2026年7月16日
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
vision-language-action
embodied-navigation
VLN-CE
ObjectNav
EQA
human-following
video-VLM
task-unification
Vicuna-7B
RSS2025
2026年7月16日
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
benchmark
mllm-agent
embodied-ai
ai2-thor
habitat
vlmbench
alfred
vision-language-action
capability-oriented-evaluation
icml-2025
2026年7月16日
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
survey
vla
vision-language-action
robotic-manipulation
taxonomy
monolithic-models
hierarchical-models
world-models
reinforcement-learning
embodied-ai
2026年7月16日
Survey of Vision-Language-Action Models for Embodied Manipulation
survey
vision-language-action
embodied-manipulation
robot-learning
imitation-learning
reinforcement-learning
action-tokenization
chain-of-thought
hierarchical-vla
benchmark
2026年7月16日
VLA 谱系(RT-1→RT-2→OpenVLA→Octo→π0→GR00T→Gemini Robotics)
vla
vision-language-action
generalist-policy
action-head
flow-matching
diffusion-policy
dual-system
cross-embodiment
autoregressive-action
world-model
deep-dive
2026年7月16日
Scaling Instructable Agents Across Many Simulated Worlds
world-model
embodied-agent
instruction-following
vision-language-action
keyboard-mouse-control
behavioral-cloning
classifier-free-guidance
video-games
sima
google-deepmind
2026年7月16日
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision–Language–Action for Autonomous Driving
driving-world-model
vision-language-action
latent-space-world-model
flow-matching
diffusion-transformer
bev-representation
navsim
nuscenes
closed-loop-planning
internvl