English
Related papers

Related papers: ManiSoft: Towards Vision-Language Manipulation for…

200 papers

Planning contact interactions is one of the core challenges of many robotic tasks. Optimizing contact locations while taking dynamics into account is computationally costly and, in environments that are only partially observable, executing…

Robotics · Computer Science 2020-04-20 Alina Kloss , Maria Bauza , Jiajun Wu , Joshua B. Tenenbaum , Alberto Rodriguez , Jeannette Bohg

Bimanual manipulation is a longstanding challenge in robotics due to the large number of degrees of freedom and the strict spatial and temporal synchronization required to generate meaningful behavior. Humans learn bimanual manipulation…

Robotics · Computer Science 2024-05-07 Arpit Bahety , Priyanka Mandikal , Ben Abbatematteo , Roberto Martín-Martín

Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes. However, current VLA models face significant challenges: they are slow during…

Recently, natural language has been the primary medium for human-robot interaction. However, its inherent lack of spatial precision introduces challenges for robotic task definition such as ambiguity and verbosity. Moreover, in some public…

Robotics · Computer Science 2025-07-29 Yanbang Li , Ziyang Gong , Haoyang Li , Xiaoqi Huang , Haolan Kang , Guangping Bai , Xianzheng Ma

Dexterous robotic hands are essential for performing complex manipulation tasks, yet remain difficult to train due to the challenges of demonstration collection and high-dimensional control. While reinforcement learning (RL) can alleviate…

While the integration of Multi-modal Large Language Models (MLLMs) with robotic systems has significantly improved robots' ability to understand and execute natural language instructions, their performance in manipulation tasks remains…

Robotics · Computer Science 2024-08-23 Siyuan Huang , Iaroslav Ponomarenko , Zhengkai Jiang , Xiaoqi Li , Xiaobin Hu , Peng Gao , Hongsheng Li , Hao Dong

Vision Language Models (VLMs) play a crucial role in robotic manipulation by enabling robots to understand and interpret the visual properties of objects and their surroundings, allowing them to perform manipulation based on this multimodal…

Robotics · Computer Science 2025-05-21 Nurhan Bulus Guran , Hanchi Ren , Jingjing Deng , Xianghua Xie

From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably. In spite of the recent adoption of vision language models (VLMs) to control robot…

ControlNet has enabled detailed spatial control in text-to-image diffusion models by incorporating additional visual conditions such as depth or edge maps. However, its effectiveness heavily depends on the availability of visual conditions…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Woosung Joung , Daewon Chae , Jinkyu Kim

Human-robot object handover is a crucial element for assistive robots that aim to help people in their daily lives, including elderly care, hospitals, and factory floors. The existing approaches to solving these tasks rely on pre-selected…

Robotics · Computer Science 2025-08-06 Lucas Chen , Guna Avula , Hanwen Ren , Zixing Wang , Ahmed H. Qureshi

In real-world scenarios, human dialogues are multi-round and diverse. Furthermore, human instructions can be unclear and human responses are unrestricted. Interactive robots face difficulties in understanding human intents and generating…

Robotics · Computer Science 2023-08-09 Zhe Zhang , Wei Chai , Jiankun Wang

How can we imbue robots with the ability to manipulate objects precisely but also to reason about them in terms of abstract concepts? Recent works in manipulation have shown that end-to-end networks can learn dexterous skills that require…

Robotics · Computer Science 2021-09-27 Mohit Shridhar , Lucas Manuelli , Dieter Fox

While leveraging abundant human videos and simulated robot data poses a scalable solution to the scarcity of real-world robot data, the generalization capability of existing vision-language-action models (VLAs) remains limited by mismatches…

Diffusion models have been verified to be effective in generating complex distributions from natural images to motion trajectories. Recent diffusion-based methods show impressive performance in 3D robotic manipulation tasks, whereas they…

Robotics · Computer Science 2025-09-09 Guanxing Lu , Zifeng Gao , Tianxing Chen , Wenxun Dai , Ziwei Wang , Wenbo Ding , Yansong Tang

While visuomotor policy learning has advanced robotic manipulation, precisely executing contact-rich tasks remains challenging due to the limitations of vision in reasoning about physical interactions. To address this, recent work has…

Robotics · Computer Science 2024-10-29 Venkatesh Pattabiraman , Yifeng Cao , Siddhant Haldar , Lerrel Pinto , Raunaq Bhirangi

Soft object manipulation has recently gained popularity within the robotics community due to its potential applications in many economically important areas. Although great progress has been recently achieved in these types of tasks, most…

Robotics · Computer Science 2021-10-20 Peng Zhou , Jihong Zhu , Shengzeng Huo , David Navarro-Alarcon

In this work, we aim to learn dexterous manipulation of deformable objects using multi-fingered hands. Reinforcement learning approaches for dexterous rigid object manipulation would struggle in this setting due to the complexity of physics…

Computer Vision and Pattern Recognition · Computer Science 2023-04-07 Sizhe Li , Zhiao Huang , Tao Chen , Tao Du , Hao Su , Joshua B. Tenenbaum , Chuang Gan

For robots to become efficient helpers in the home, they must learn to perform new mobile manipulation tasks simply by watching humans perform them. Learning from a single video demonstration from a human is challenging as the robot needs…

Robotics · Computer Science 2025-06-23 Arpit Bahety , Arnav Balaji , Ben Abbatematteo , Roberto Martín-Martín

Learning a generalizable bimanual manipulation policy is extremely challenging for embodied agents due to the large action space and the need for coordinated arm movements. Existing approaches rely on Vision-Language-Action (VLA) models to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Chenyou Fan , Fangzheng Yan , Chenjia Bai , Jiepeng Wang , Chi Zhang , Zhen Wang , Xuelong Li

Visual Robot Manipulation (VRM) aims to enable a robot to follow natural language instructions based on robot states and visual observations, and therefore requires costly multi-modal data. To compensate for the deficiency of robot data,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Dejie Yang , Zijing Zhao , Yang Liu
‹ Prev 1 4 5 6 7 8 10 Next ›