English
Related papers

Related papers: Grounding 3D Object Affordance with Language Instr…

200 papers

Ground robot navigation in complex 3D environments is often hindered by geometric ambiguity, where non-traversable structures such as furniture share local geometric properties with navigable ground. Furthermore, the computational cost of…

Robotics · Computer Science 2026-05-19 Yifan Du , Chengwei Zhang , Siyu Liao , Zhongfeng Wang

Furniture assembly is a crucial yet challenging task for robots, requiring precise dual-arm coordination where one arm manipulates parts while the other provides collaborative support and stabilization. To accomplish this task more…

Robotics · Computer Science 2026-01-19 Jiaqi Liang , Yue Chen , Qize Yu , Yan Shen , Haipeng Zhang , Hao Dong , Ruihai Wu

Accurate affordance detection and segmentation with pixel precision is an important piece in many complex systems based on interactions, such as robots and assitive devices. We present a new approach to affordance perception which enables…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Lorenzo Mur-Labadia , Jose J. Guerrero , Ruben Martinez-Cantin

Robots operating in human-centered environments should have the ability to understand how objects function: what can be done with each object, where this interaction may occur, and how the object is used to achieve a goal. To this end, we…

Perception-for-grasping is a challenging problem in robotics. Inexpensive range sensors such as the Microsoft Kinect provide sensing capabilities that have given new life to the effort of developing robust and accurate perception methods…

Robotics · Computer Science 2013-11-14 Andreas ten Pas , Robert Platt

Dexterous robotic hands are appealing for their agility and human-like morphology, yet their high degree of freedom makes learning to manipulate challenging. We introduce an approach for learning dexterous grasping. Our key idea is to embed…

Robotics · Computer Science 2021-06-18 Priyanka Mandikal , Kristen Grauman

Grounding object properties and relations in 3D scenes is a prerequisite for a wide range of artificial intelligence tasks, such as visually grounded dialogues and embodied manipulation. However, the variability of the 3D domain induces two…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Joy Hsu , Jiayuan Mao , Jiajun Wu

Learning object affordances is an effective tool in the field of robot learning. While the data-driven models investigate affordances of single or paired objects, there is a gap in the exploration of affordances of compound objects composed…

Robotics · Computer Science 2024-12-18 Tuba Girgin , Emre Ugur

Precise object placement remains underexplored in aerial manipulation, where most systems rely on predefined target coordinates and focus primarily on grasping and control. Specifying exact placement poses, however, is cumbersome in…

Robotics · Computer Science 2026-03-10 Sarthak Mishra , Rishabh Dev Yadav , Naveen Nair , Wei Pan , Spandan Roy

Language-guided active sensing is a robotics subtask where a robot with an onboard sensor interacts efficiently with the environment via object manipulation to maximize perceptual information, following given language instructions. These…

Robotics · Computer Science 2024-02-06 Weihan Chen , Hanwen Ren , Ahmed H. Qureshi

Most models tasked to ground referential utterances in 2D and 3D scenes learn to select the referred object from a pool of object proposals provided by a pre-trained detector. This is limiting because an utterance may refer to visual…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Ayush Jain , Nikolaos Gkanatsios , Ishita Mediratta , Katerina Fragkiadaki

The choice of a grasp plays a critical role in the success of downstream manipulation tasks. Consider a task of placing an object in a cluttered scene; the majority of possible grasps may not be suitable for the desired placement. In this…

Robotics · Computer Science 2023-04-11 Zhanpeng He , Nikhil Chavan-Dafle , Jinwook Huh , Shuran Song , Volkan Isler

Multi-object 3D Grounding involves locating 3D boxes based on a given query phrase from a point cloud. It is a challenging and significant task with numerous applications in visual understanding, human-computer interaction, and robotics. To…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Haomeng Zhang , Chiao-An Yang , Raymond A. Yeh

Visual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, single-platform constraints, and small scale. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Rong Li , Yuhao Dong , Tianshuai Hu , Ao Liang , Youquan Liu , Dongyue Lu , Liang Pan , Lingdong Kong , Junwei Liang , Ziwei Liu

To be capable of lifelong learning in a real-life environment, robots have to tackle multiple challenges. Being able to relate physical properties they may observe in their environment to possible interactions they may have is one of them.…

Artificial Intelligence · Computer Science 2020-09-24 Alexandre Manoury , Sao Mai Nguyen , Cédric Buche

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Ziyu Zhu , Xilin Wang , Yixuan Li , Zhuofan Zhang , Xiaojian Ma , Yixin Chen , Baoxiong Jia , Wei Liang , Qian Yu , Zhidong Deng , Siyuan Huang , Qing Li

3D visual grounding aims to identify the target object within a 3D point cloud scene referred to by a natural language description. Previous works usually require significant data relating to point color and their descriptions to exploit…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Tung-Yu Wu , Sheng-Yu Huang , Yu-Chiang Frank Wang

Grasp detection in clutter requires the robot to reason about the 3D scene from incomplete and noisy perception. In this work, we draw insight that 3D reconstruction and grasp learning are two intimately connected tasks, both of which…

Robotics · Computer Science 2021-07-22 Zhenyu Jiang , Yifeng Zhu , Maxwell Svetlik , Kuan Fang , Yuke Zhu

Nowadays, robots are dominating the manufacturing, entertainment and healthcare industries. Robot vision aims to equip robots with the ability to discover information, understand it and interact with the environment. These capabilities…

Computer Vision and Pattern Recognition · Computer Science 2018-07-19 Mohammed Hassanin , Salman Khan , Murat Tahtali

Controlling robots to perform tasks via natural language is one of the most challenging topics in human-robot interaction. In this work, we present a robot system that follows unconstrained language instructions to pick and place arbitrary…

Robotics · Computer Science 2021-02-17 Oier Mees , Wolfram Burgard