中文
相关论文

相关论文: FALCON: Actively Decoupled Visuomotor Policies for…

200 篇论文

Humanoid loco-manipulation holds transformative potential for daily service and industrial tasks, yet achieving precise, robust whole-body control with 3D end-effector force interaction remains a major challenge. Prior approaches are often…

Diffusion policies are widely adopted in complex visuomotor tasks for their ability to capture multimodal action distributions. However, the multiple sampling steps required for action generation significantly harm real-time inference…

Visuomotor policies aim to learn complex manipulation tasks from expert demonstrations. However, generating smooth and coherent trajectories remains challenging, as it requires balancing proximal precision with distal foresight. Existing…

机器人学 · 计算机科学 2026-05-21 Qian He , Zhenshuo Yang , Wenqi Liang , Chunhui Hao , Nicu Sebe , Jiandong Tian

We introduce FALCON, a unified self-supervised video pretraining approach for UAV action recognition from raw RGB aerial footage, requiring no additional preprocessing at inference. UAV videos exhibit severe spatial imbalance: large,…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Ruiqi Xian , Xiyang Wu , Tianrui Guan , Xijun Wang , Boqing Gong , Dinesh Manocha

Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require…

The incorporation of high-resolution visual input equips multimodal large language models (MLLMs) with enhanced visual perception capabilities for real-world tasks. However, most existing high-resolution MLLMs rely on a cropping-based…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Renshan Zhang , Rui Shao , Gongwei Chen , Miao Zhang , Kaiwen Zhou , Weili Guan , Liqiang Nie

While recent advances have demonstrated strong performance in individual humanoid skills such as upright locomotion, fall recovery and whole-body coordination, learning a single policy that masters all these skills remains challenging due…

机器人学 · 计算机科学 2026-03-05 Dewei Wang , Xinmiao Wang , Chenyun Zhang , Jiyuan Shi , Yingnan Zhao , Chenjia Bai , Xuelong Li

Recent advances in legged locomotion learning are still dominated by the utilization of geometric representations of the environment, limiting the robot's capability to respond to higher-level semantics such as human instructions. To…

机器人学 · 计算机科学 2026-02-12 I Made Aswin Nahrendra , Seunghyun Lee , Dongkyu Lee , Hyun Myung

This work introduces DiffuseLoco, a framework for training multi-skill diffusion-based policies for dynamic legged locomotion from offline datasets, enabling real-time control of diverse skills on robots in the real world. Offline learning…

机器人学 · 计算机科学 2024-05-01 Xiaoyu Huang , Yufeng Chi , Ruofeng Wang , Zhongyu Li , Xue Bin Peng , Sophia Shao , Borivoje Nikolic , Koushil Sreenath

Many recent Vision-Language-Action models employ diffusion or flow-matching backbones with hundreds of millions of parameters for action generation. However, unlike image synthesis where the output spans millions of diverse pixels, a…

机器人学 · 计算机科学 2026-03-17 Jian Zhou , Sihao Lin , Shuai Fu , Zerui Li , Gengze Zhou , Qi WU

Opening heavy, self closing doors, especially those that require pulling remains a long standing challenge in robotics. Humans naturally employ both arms in a dexterous manner, rotating the handle, widening the gap, holding the door,…

机器人学 · 计算机科学 2026-05-18 Shangqun Yu , Matthew En , Daniel Wu , Sangjun Park , Ziyi Zhou , Seyed Fakoorian , Donghyun Kim

While recent large vision-language models (VLMs) have improved generalization in vision-language navigation (VLN), existing methods typically rely on end-to-end pipelines that map vision-language inputs directly to short-horizon discrete…

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of…

机器人学 · 计算机科学 2024-02-06 Xinghang Li , Minghuan Liu , Hanbo Zhang , Cunjun Yu , Jie Xu , Hongtao Wu , Chilam Cheang , Ya Jing , Weinan Zhang , Huaping Liu , Hang Li , Tao Kong

Flow-matching-based policies have recently emerged as a promising approach for learning-based robot manipulation, offering significant acceleration in action sampling compared to diffusion-based policies. However, conventional flow-matching…

机器人学 · 计算机科学 2025-10-03 Xuanran Zhai , Qianyou Zhao , Qiaojun Yu , Ce Hao

Generalizing locomotion policies across diverse legged robots with varying morphologies is a key challenge due to differences in observation/action dimensions and system dynamics. In this work, we propose Multi-Loco, a novel unified…

机器人学 · 计算机科学 2025-06-16 Shunpeng Yang , Zhen Fu , Zhefeng Cao , Guo Junde , Patrick Wensing , Wei Zhang , Hua Chen

Most Vision-Language-Action (VLA) systems integrate a Vision-Language Model (VLM) for semantic reasoning with an action expert generating continuous action signals, yet both typically run at a single unified frequency. As a result, policy…

机器人学 · 计算机科学 2025-12-24 Teqiang Zou , Hongliang Zeng , Yuxuan Nong , Yifan Li , Kehui Liu , Haotian Yang , Xinyang Ling , Xin Li , Lianyang Ma

In this study, we address vision-language-guided multi-robot cooperative transport, where each robot grounds natural-language instructions from onboard camera observations. A key challenge in this decentralized setting is perceptual…

机器人学 · 计算机科学 2026-02-10 Joachim Yann Despature , Kazuki Shibata , Takamitsu Matsubara

Diffusion policies have demonstrated strong performance in generative modeling, making them promising for robotic manipulation guided by natural language instructions. However, generalizing language-conditioned diffusion policies to…

机器人学 · 计算机科学 2025-08-20 Ce Hao , Kelvin Lin , Zhiwei Xue , Siyuan Luo , Harold Soh

Current multi-modal models exhibit a notable misalignment with the human visual system when identifying objects that are visually assimilated into the background. Our observations reveal that these multi-modal models cannot distinguish…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Ruolin Shen , Xiaozhong Ji , Kai WU , Jiangning Zhang , Yijun He , HaiHua Yang , Xiaobin Hu , Xiaoyu Sun

Visuomotor imitation learning policies enable robots to efficiently acquire manipulation skills from visual demonstrations. However, as scene complexity and visual distractions increase, policies that perform well in simple settings often…

‹ 上一页 1 2 3 10 下一页 ›