中文
相关论文

相关论文: CompassAD: Intent-Driven 3D Affordance Grounding i…

200 篇论文

Intelligent embodied agents should not simply follow instructions, as real-world environments often involve unexpected conditions and exceptions. However, existing methods usually focus on directly executing instructions, without…

人工智能 · 计算机科学 2026-04-21 Pei-An Chen , Yong-Ching Liang , Jia-Fong Yeh , Hung-Ting Su , Yi-Ting Chen , Min Sun , Winston Hsu

For robots to exhibit a high level of intelligence in the real world, they must be able to assess objects for which they have no prior knowledge. Therefore, it is crucial for robots to perceive object affordances by reasoning about physical…

机器人学 · 计算机科学 2020-04-09 Hongtao Wu , Deven Misra , Gregory S. Chirikjian

3D object affordance grounding aims to predict the touchable regions on a 3d object, which is crucial for human-object interaction, human-robot interaction, embodied perception, and robot learning. Recent advances tackle this problem via…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Hanqing Wang , Zhenhao Zhang , Kaiyang Ji , Mingyu Liu , Wenti Yin , Yuchao Chen , Zhirui Liu , Xiangyu Zeng , Tianxiang Gui , Hangxing Zhang

Affordance grounding refers to the task of finding the area of an object with which one can interact. It is a fundamental but challenging task, as a successful solution requires the comprehensive understanding of a scene in multiple aspects…

计算机视觉与模式识别 · 计算机科学 2024-04-19 Shengyi Qian , Weifeng Chen , Min Bai , Xiong Zhou , Zhuowen Tu , Li Erran Li

Robotic manipulation with two-finger grippers is challenged by objects lacking distinct graspable features. Traditional pre-grasping methods, which typically involve repositioning objects or utilizing external aids like table edges, are…

机器人学 · 计算机科学 2024-08-26 Kairui Ding , Boyuan Chen , Ruihai Wu , Yuyang Li , Zongzheng Zhang , Huan-ang Gao , Siqi Li , Guyue Zhou , Yixin Zhu , Hao Dong , Hao Zhao

Recent advances in generative AI have dramatically improved photorealistic image synthesis, yet they fall short for studio-level multi-object compositing. This task demands simultaneous (i) near-perfect preservation of each item's identity,…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Gemma Canet Tarrés , Manel Baradad , Francesc Moreno-Noguer , Yumeng Li

Inferring affordable (i.e., graspable) parts of arbitrary objects based on human specifications is essential for robots advancing toward open-vocabulary manipulation. Current grasp planners, however, are hindered by limited vision-language…

机器人学 · 计算机科学 2025-05-02 Teli Ma , Zifan Wang , Jiaming Zhou , Mengmeng Wang , Junwei Liang

Identifying affordance regions on 3D objects from semantic cues is essential for robotics and human-machine interaction. However, existing 3D affordance learning methods struggle with generalization and robustness due to limited annotated…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Dongyue Lu , Lingdong Kong , Tianxin Huang , Gim Hee Lee

6D object pose estimation has been a research topic in the field of computer vision and robotics. Many modern world applications like robot grasping, manipulation, autonomous navigation etc, require the correct pose of objects present in a…

计算机视觉与模式识别 · 计算机科学 2022-12-13 Ankit Kumar , Priya Shukla , Vandana Kushwaha , G. C. Nandi

Affordance information about a scene provides important clues as to what actions may be executed in pursuit of meeting a specified goal state. Thus, integrating affordance-based reasoning into symbolic action plannning pipelines would…

机器人学 · 计算机科学 2020-09-15 Fu-Jen Chu , Ruinian Xu , Chao Tang , Patricio A. Vela

Vision-based robot learning often relies on dense image or point-cloud inputs, which are computationally heavy and entangle irrelevant background features. Existing keypoint-based approaches can focus on manipulation-centric features and be…

机器人学 · 计算机科学 2026-04-17 Anukriti Singh , Kasra Torshizi , Khuzema Habib , Kelin Yu , Ruohan Gao , Pratap Tokekar

In real-life scenarios, humans seek out objects in the 3D world to fulfill their daily needs or intentions. This inspires us to introduce 3D intention grounding, a new task in 3D object detection employing RGB-D, based on human intention,…

计算机视觉与模式识别 · 计算机科学 2025-02-19 Weitai Kang , Mengxue Qu , Jyoti Kini , Yunchao Wei , Mubarak Shah , Yan Yan

Translating high-level linguistic instructions into precise robotic actions in the physical world remains challenging, particularly when considering the feasibility of interacting with 3D objects. In this paper, we introduce 3D-TAFS, a…

机器人学 · 计算机科学 2025-04-08 Meng Chu , Xuan Zhang , Zhedong Zheng , Tat-Seng Chua

Grounded understanding of natural language in physical scenes can greatly benefit robots that follow human instructions. In object manipulation scenarios, existing end-to-end models are proficient at understanding semantic concepts, but…

机器人学 · 计算机科学 2023-04-03 Qian Luo , Yunfei Li , Yi Wu

Humans excel at acquiring knowledge through observation. For example, we can learn to use new tools by watching demonstrations. This skill is fundamental for intelligent systems to interact with the world. A key step to acquire this skill…

计算机视觉与模式识别 · 计算机科学 2023-03-20 Gen Li , Varun Jampani , Deqing Sun , Laura Sevilla-Lara

Camouflaged Object Detection (COD) aims to segment targets that share extreme textural and structural similarities with their complex environments. Leveraging their capacity for long-range dependency modeling, Transformer-based detectors…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Yuhan Gao , Shuhao Kang , Xin He , Bing Li , Xu Cheng , Yun Liu

Recent advancements in 3D perception systems have significantly improved their ability to perform visual recognition tasks such as segmentation. However, these systems still heavily rely on explicit human instruction to identify target…

计算机视觉与模式识别 · 计算机科学 2024-04-08 Amrin Kareem , Jean Lahoud , Hisham Cholakkal

The physical and textural attributes of objects have been widely studied for recognition, detection and segmentation tasks in computer vision.~A number of datasets, such as large scale ImageNet, have been proposed for feature learning using…

计算机视觉与模式识别 · 计算机科学 2023-09-14 Zeyad Khalifa , Syed Afaq Ali Shah

3D visual grounding involves finding a target object in a 3D scene that corresponds to a given sentence query. Although many approaches have been proposed and achieved impressive performance, they all require dense object-sentence pair…

计算机视觉与模式识别 · 计算机科学 2023-07-19 Zehan Wang , Haifeng Huang , Yang Zhao , Linjun Li , Xize Cheng , Yichen Zhu , Aoxiong Yin , Zhou Zhao

Acquiring knowledge about object interactions and affordances can facilitate scene understanding and human-robot collaboration tasks. As humans tend to use objects in many different ways depending on the scene and the objects' availability,…

人工智能 · 计算机科学 2023-04-13 Alexia Toumpa , Anthony G. Cohn