中文
相关论文

相关论文: Learning to Act with Affordance-Aware Multimodal N…

200 篇论文

We address the problem of affordance reasoning in diverse scenes that appear in the real world. Affordances relate the agent's actions to their effects when taken on the surrounding objects. In our work, we take the egocentric view of the…

计算机视觉与模式识别 · 计算机科学 2018-06-18 Ching-Yao Chuang , Jiaman Li , Antonio Torralba , Sanja Fidler

Planning is a fundamental task in artificial intelligence that involves finding a sequence of actions that achieve a specified goal in a given environment. Large language models (LLMs) are increasingly used for applications that require…

计算与语言 · 计算机科学 2024-05-24 Eran Hirsch , Guy Uziel , Ateret Anaby-Tavor

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single-affordance paradigms, where each affordance type or…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Chunlin Yu , Hanqing Wang , Ye Shi , Haoyang Luo , Sibei Yang , Jingyi Yu , Jingya Wang

Simultaneous localization and mapping (SLAM) is a foundational state estimation problem in robotics in which a robot accurately constructs a map of its environment while also localizing itself within this construction. We study the active…

机器人学 · 计算机科学 2026-04-24 Ilir Gusija , Fady Alajaji , Serdar Yüksel

Neural implicit representations have had a significant impact on simultaneous localization and mapping (SLAM) by enabling robots to build continuous, differentiable, and high-fidelity 3D maps from sensor data. However, as the scale and…

机器人学 · 计算机科学 2025-04-29 Yulun Tian , Hanwen Cao , Sunghwan Kim , Nikolay Atanasov

Grounding the reasoning ability of large language models (LLMs) for embodied tasks is challenging due to the complexity of the physical world. Especially, LLM planning for multi-agent collaboration requires communication of agents or credit…

人工智能 · 计算机科学 2025-09-30 Yang Zhang , Shixin Yang , Chenjia Bai , Fei Wu , Xiu Li , Zhen Wang , Xuelong Li

Visual navigation for autonomous agents is a core task in the fields of computer vision and robotics. Learning-based methods, such as deep reinforcement learning, have the potential to outperform the classical solutions developed for this…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Zachary Seymour , Kowshik Thopalli , Niluthpol Mithun , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

Aerial navigation in GPS-denied, indoor environments, is still an open challenge. Drones can perceive the environment from a richer set of viewpoints, while having more stringent compute and energy constraints than other autonomous…

机器人学 · 计算机科学 2021-06-18 Ni Wang , Ozan Catal , Tim Verbelen , Matthias Hartmann , Bart Dhoedt

We present ESLAM, an efficient implicit neural representation method for Simultaneous Localization and Mapping (SLAM). ESLAM reads RGB-D frames with unknown camera poses in a sequential manner and incrementally reconstructs the scene…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Mohammad Mahdi Johari , Camilla Carta , François Fleuret

Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We adopt a task-oriented perspective to systematically review the applications and advancements of multimodal fusion…

While large language models (LLMs) are successful in completing various language processing tasks, they easily fail to interact with the physical world by generating control sequences properly. We find that the main reason is that LLMs are…

人工智能 · 计算机科学 2024-04-18 Guangran Cheng , Chuheng Zhang , Wenzhe Cai , Li Zhao , Changyin Sun , Jiang Bian

Imitation learning has unlocked the potential for robots to exhibit highly dexterous behaviours. However, it still struggles with long-horizon, multi-object tasks due to poor sample efficiency and limited generalisation. Existing methods…

机器人学 · 计算机科学 2025-09-05 Krishan Rana , Jad Abou-Chakra , Sourav Garg , Robert Lee , Ian Reid , Niko Suenderhauf

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Dongming Wu , Yanping Fu , Saike Huang , Yingfei Liu , Fan Jia , Nian Liu , Feng Dai , Tiancai Wang , Rao Muhammad Anwer , Fahad Shahbaz Khan , Jianbing Shen

Recent advances in Multimodal Large Language Models (MLLMs) have shown impressive reasoning capabilities across vision-language tasks, yet still face the challenge of compute-difficulty mismatch. Through empirical analyses, we identify that…

机器学习 · 计算机科学 2026-03-17 Huijie Guo , Jingyao Wang , Lingyu Si , Jiahuan Zhou , Changwen Zheng , Wenwen Qiang

Simultaneous Localization And Mapping (SLAM) is a task to estimate the robot location and to reconstruct the environment based on observation from sensors such as LIght Detection And Ranging (LiDAR) and camera. It is widely used in robotic…

机器人学 · 计算机科学 2021-02-18 Han Wang , Chen Wang , Lihua Xie

A fundamental challenge in robot navigation lies in learning policies that generalize across diverse environments while conforming to the unique physical constraints and capabilities of a specific embodiment (e.g., quadrupeds can walk up…

Edge computing is increasingly proposed as a solution for reducing resource consumption of mobile devices running simultaneous localization and mapping (SLAM) algorithms, with most edge-assisted SLAM systems assuming the communication…

系统与控制 · 电气工程与系统科学 2023-01-12 Ying Chen , Hazer Inaltekin , Maria Gorlatova

Visual understanding of 3D environments in real-time, at low power, is a huge computational challenge. Often referred to as SLAM (Simultaneous Localisation and Mapping), it is central to applications spanning domestic and industrial…

Large-scale pretrained foundation models have been an emerging paradigm for building artificial intelligence (AI) systems, which can be quickly adapted to a wide range of downstream tasks. This paper presents mPLUG, a new vision-language…

Topological strategies for navigation meaningfully reduce the space of possible actions available to a robot, allowing use of heuristic priors or learning to enable computationally efficient, intelligent planning. The challenges in…

机器人学 · 计算机科学 2020-04-01 Gregory J. Stein , Christopher Bradley , Victoria Preston , Nicholas Roy