English
Related papers

Related papers: VoxAct-B: Voxel-Based Acting and Stabilizing Polic…

200 papers

We present a method for learning a human-robot collaboration policy from human-human collaboration demonstrations. An effective robot assistant must learn to handle diverse human behaviors shown in the demonstrations and be robust when the…

Robotics · Computer Science 2023-09-21 Chen Wang , Claudia Pérez-D'Arpino , Danfei Xu , Li Fei-Fei , C. Karen Liu , Silvio Savarese

This work developed collaborative bimanual manipulation for reliable and safe human-robot collaboration, which allows remote and local human operators to work interactively for bimanual tasks. We proposed an optimal motion adaptation to…

Robotics · Computer Science 2023-07-19 Ruoshi Wen , Quentin Rouxel , Michael Mistry , Zhibin Li , Carlo Tiseo

Generating large-scale demonstrations for dexterous hand manipulation remains challenging, and several approaches have been proposed in recent years to address this. Among them, generative models have emerged as a promising paradigm,…

Robotics · Computer Science 2025-06-23 Jianglong Ye , Keyi Wang , Chengjing Yuan , Ruihan Yang , Yiquan Li , Jiyue Zhu , Yuzhe Qin , Xueyan Zou , Xiaolong Wang

Learning natural, stable, and compositionally generalizable whole-body control policies for humanoid robots performing simultaneous locomotion and manipulation (loco-manipulation) remains a fundamental challenge in robotics. Existing…

Controlling robots through natural language is pivotal for enhancing human-robot collaboration and synthesizing complex robot behaviors. Recent works that are trained on large robot datasets show impressive generalization abilities.…

3D hand shape and pose estimation from a single depth map is a new and challenging computer vision problem with many applications. The state-of-the-art methods directly regress 3D hand meshes from 2D depth images via 2D convolutional neural…

Computer Vision and Pattern Recognition · Computer Science 2020-04-06 Jameel Malik , Ibrahim Abdelaziz , Ahmed Elhayek , Soshi Shimada , Sk Aziz Ali , Vladislav Golyanik , Christian Theobalt , Didier Stricker

Goal-oriented planning, or anticipating a series of actions that transition an agent from its current state to a predefined objective, is crucial for developing intelligent assistants aiding users in daily procedural tasks. The problem…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Md Mohaiminul Islam , Tushar Nagarajan , Huiyu Wang , Fu-Jen Chu , Kris Kitani , Gedas Bertasius , Xitong Yang

Most existing vision-language manipulation research targets rigid robotic arms, whose fixed morphology limits adaptability in cluttered or confined spaces. Soft robotic arms offer an appealing alternative due to their deformability, but…

Robotics · Computer Science 2026-05-19 Ziyu Wei , Luting Wang , Chen Gao , Li Wen , Si Liu

Foundation models applied in robotics, particularly \textbf{Vision--Language--Action (VLA)} models, hold great promise for achieving general-purpose manipulation. Yet, systematic real-world evaluations and cross-model comparisons remain…

Robotics · Computer Science 2025-11-17 Yihao Zhang , Yuankai Qi , Xi Zheng

To perform household tasks, assistive robots receive commands in the form of user language instructions for tool manipulation. The initial stage involves selecting the intended tool (i.e., object grounding) and grasping it in a…

Robotics · Computer Science 2023-03-01 Chao Tang , Dehao Huang , Lingxiao Meng , Weiyu Liu , Hong Zhang

Robotic systems that aspire to operate in uninstrumented real-world environments must perceive the world directly via onboard sensing. Vision-based learning systems aim to eliminate the need for environment instrumentation by building an…

Robotics · Computer Science 2024-05-14 Patrick Lancaster , Nicklas Hansen , Aravind Rajeswaran , Vikash Kumar

Multi-finger robotic hand manipulation and grasping are challenging due to the high-dimensional action space and the difficulty of acquiring large-scale training data. Existing approaches largely rely on human teleoperation with wearable…

Learning robot manipulation from human videos is appealing due to the scale and diversity of human demonstrations, but transferring such demonstrations to executable robot behavior remains challenging. Prior work either relies on robot data…

Robotics · Computer Science 2026-05-05 Yifan Han , Jianxiang Liu , Haoyu Zhang , Yuqi Gu , Yunhan Guo , Wenzhao Lian

We present ArtiGrasp, a novel method to synthesize bi-manual hand-object interactions that include grasping and articulation. This task is challenging due to the diversity of the global wrist motions and the precise finger control that are…

Robotics · Computer Science 2024-03-05 Hui Zhang , Sammy Christen , Zicong Fan , Luocheng Zheng , Jemin Hwangbo , Jie Song , Otmar Hilliges

Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Chi-Pin Huang , Yueh-Hua Wu , Min-Hung Chen , Yu-Chiang Frank Wang , Fu-En Yang

When performing 3D manipulation tasks, robots have to execute action planning based on perceptions from multiple fixed cameras. The multi-camera setup introduces substantial redundancy and irrelevant information, which increases…

Robotics · Computer Science 2025-12-19 Yixiang Chen , Yan Huang , Keji He , Peiyan Li , Liang Wang

Long-horizon contact-rich robotic manipulation remains challenging due to partial observability and unstable subtask transitions under contact uncertainty. While hierarchical architectures improve temporal reasoning and bilateral imitation…

Robotics · Computer Science 2026-03-27 Thanpimon Buamanee , Masato Kobayashi , Yuki Uranishi

Bimanual manipulation tasks typically involve multiple stages which require efficient interactions between two arms, posing step-wise and stage-wise challenges for imitation learning systems. Specifically, failure and delay of one step will…

Robotics · Computer Science 2024-09-05 Dongjie Yu , Hang Xu , Yizhou Chen , Yi Ren , Jia Pan

Learning generalizable policies for robotic manipulation increasingly relies on large-scale models that map language instructions to actions (L2A). However, this one-way paradigm often produces policies that execute tasks without deeper…

Robotics · Computer Science 2026-05-25 Youngjin Hong , Houjian Yu , Mingen Li , Changhyun Choi

This paper presents a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot…