English
Related papers

Related papers: AffordBot: 3D Fine-grained Embodied Reasoning via …

200 papers

Artificial intelligence is essential to succeed in challenging activities that involve dynamic environments, such as object manipulation tasks in indoor scenes. Most of the state-of-the-art literature explores robotic grasping methods by…

Robotics · Computer Science 2019-05-28 Paola Ardón , Èric Pairet , Ron Petrick , Subramanian Ramamoorthy , Katrin Lohan

3D object grounding localizes referred objects in a 3D scene from natural language. Unified instance-centric 3D-LLMs aim to solve grounding together with dialog, QA, and captioning, yet many rely on a single pointer-style grounding decision…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Jiawei Li , Ziyi Liu , Weijie Shi , Long Chen , Jiajie Xu , Xiaofang Zhou

Object affordance is an important concept in hand-object interaction, providing information on action possibilities based on human motor capacity and objects' physical property thus benefiting tasks such as action anticipation and robot…

Computer Vision and Pattern Recognition · Computer Science 2023-02-13 Zecheng Yu , Yifei Huang , Ryosuke Furuta , Takuma Yagi , Yusuke Goutsu , Yoichi Sato

The remarkable potential of multi-modal large language models (MLLMs) in comprehending both vision and language information has been widely acknowledged. However, the scarcity of 3D scenes-language pairs in comparison to their 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Zeju Li , Chao Zhang , Xiaoyan Wang , Ruilong Ren , Yifan Xu , Ruifei Ma , Xiangde Liu

Robotic agents need to understand how to interact with objects in their environment, both autonomously and during human-robot interactions. Affordance detection on 3D point clouds, which identifies object regions that allow specific…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Maximilian Xiling Li , Korbinian Rudolf , Nils Blank , Rudolf Lioutikov

Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each step. This reasoning unfaithfulness frequently manifests as…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Chuhan Wang , Xintong Li , Jennifer Yuntong Zhang , Junda Wu , Chengkai Huang , Lina Yao , Julian McAuley , Jingbo Shang

Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents…

Open-vocabulary 3D affordance detection requires localizing interaction regions on point clouds given novel affordance descriptions. Recent methods extend multimodal large language models (MLLMs) with special output tokens that are decoded…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Haowen Sun , Shaolong Zhang , Mingyang Li , Chengzhong Ma , Xinzhe Chen , Qiongjie Cui , Xingyu Chen , Zeyang Liu , Xuguang Lan

Affordances represent the inherent effect and action possibilities that objects offer to the agents within a given context. From a theoretical viewpoint, affordances bridge the gap between effect and action, providing a functional…

Robotics · Computer Science 2024-10-11 Hakan Aktas , Yukie Nagai , Minoru Asada , Matteo Saveriano , Erhan Oztop , Emre Ugur

Humans build viewpoint-independent cognitive maps through navigation, enabling intuitive reasoning about object permanence and spatial relations. We argue that multimodal large language models (MLLMs), despite extensive video training, lack…

Machine Learning · Computer Science 2025-12-02 Jacob Thompson , Emiliano Garcia-Lopez , Yonatan Bisk

Affordance modeling plays an important role in visual understanding. In this paper, we aim to predict affordances of 3D indoor scenes, specifically what human poses are afforded by a given indoor environment, such as sitting on a chair or…

Computer Vision and Pattern Recognition · Computer Science 2019-03-19 Xueting Li , Sifei Liu , Kihwan Kim , Xiaolong Wang , Ming-Hsuan Yang , Jan Kautz

Multimodal Large Language Models (MLLMs) have made remarkable progress in multimodal perception and reasoning by bridging vision and language. However, most existing MLLMs perform reasoning primarily with textual CoT, which limits their…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Jintao Tong , Shilin Yan , Hongwei Xue , Xiaojun Tang , Kunyu Shi , Guannan Zhang , Ruixuan Li , Yixiong Zou

While large language models (LMs) have shown remarkable capabilities across numerous tasks, they often struggle with simple reasoning and planning in physical environments, such as understanding object permanence or planning household…

Computation and Language · Computer Science 2023-10-31 Jiannan Xiang , Tianhua Tao , Yi Gu , Tianmin Shu , Zirui Wang , Zichao Yang , Zhiting Hu

Building models that can understand and reason about 3D scenes is difficult owing to the lack of data sources for 3D supervised training and large-scale training regimes. In this work we ask - How can the knowledge in a pre-trained language…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Shivam Chandhok

Recent advancements in 3D perception systems have significantly improved their ability to perform visual recognition tasks such as segmentation. However, these systems still heavily rely on explicit human instruction to identify target…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Amrin Kareem , Jean Lahoud , Hisham Cholakkal

Autonomous artificial intelligence (AI) agents have emerged as promising protocols for automatically understanding the language-based environment, particularly with the exponential development of large language models (LLMs). However, a…

Computation and Language · Computer Science 2024-06-07 Jiahuan Pei , Irene Viola , Haochen Huang , Junxiao Wang , Moonisa Ahsan , Fanghua Ye , Jiang Yiming , Yao Sai , Di Wang , Zhumin Chen , Pengjie Ren , Pablo Cesar

We present a Collaborative Agent-Based Framework for Multi-Image Reasoning. Our approach tackles the challenge of interleaved multimodal reasoning across diverse datasets and task formats by employing a dual-agent system: a language-based…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Angelos Vlachos , Giorgos Filandrianos , Maria Lymperaiou , Nikolaos Spanos , Ilias Mitsouras , Vasileios Karampinis , Athanasios Voulodimos

Interactions with articulated objects are a challenging but important task for mobile robots. To tackle this challenge, we propose a novel closed-loop control pipeline, which integrates manipulation priors from affordance estimation with…

Robotics · Computer Science 2023-02-07 Giulio Schiavi , Paula Wulkop , Giuseppe Rizzi , Lionel Ott , Roland Siegwart , Jen Jen Chung

Many everyday robot manipulation skills are affordance-dependent, with success determined by whether the robot contacts the functional object region required by the subsequent action. Current simulation data generators obtain contacts from…

An embodied agent assisting humans is often asked to complete new tasks, and there may not be sufficient time or labeled examples to train the agent to perform these new tasks. Large Language Models (LLMs) trained on considerable knowledge…

‹ Prev 1 8 9 10 Next ›