中文
相关论文

相关论文: CoNav: Collaborative Cross-Modal Reasoning for Emb…

200 篇论文

While multimodal large language models (MLLMs) have made groundbreaking progress in embodied intelligence, they still face significant challenges in spatial reasoning for complex long-horizon tasks. To address this gap, we propose…

Vision-Language Models (VLMs) have made significant strides in static image understanding but continue to face critical hurdles in spatiotemporal reasoning. A major bottleneck is "multi-image reasoning hallucination", where a massive…

Collaborative 3D object detection exploits information exchange among multiple agents to enhance accuracy of object detection in presence of sensor impairments such as occlusion. However, in practice, pose estimation errors due to imperfect…

计算机视觉与模式识别 · 计算机科学 2023-03-06 Yifan Lu , Quanhao Li , Baoan Liu , Mehrdad Dianati , Chen Feng , Siheng Chen , Yanfeng Wang

Communication between embodied AI agents has received increasing attention in recent years. Despite its use, it is still unclear whether the learned communication is interpretable and grounded in perception. To study the grounding of…

计算机视觉与模式识别 · 计算机科学 2021-10-13 Shivansh Patel , Saim Wani , Unnat Jain , Alexander Schwing , Svetlana Lazebnik , Manolis Savva , Angel X. Chang

Existing language-driven embodied navigation paradigms face challenges in functional buildings (FBs) with highly similar features, as they lack the ability to effectively utilize priori spatial knowledge. To tackle this issue, we propose a…

机器人学 · 计算机科学 2026-03-11 Jiang Gao , Xiangyu Dong , Haozhou Li , Haoran Zhao , Yaoming Zhou , Xiaoguang Ma

Image-based single-modality compression learning approaches have demonstrated exceptionally powerful encoding and decoding capabilities in the past few years , but suffer from blur and severe semantics loss at extremely low bitrates. To…

图像与视频处理 · 电气工程与系统科学 2023-04-27 Xuhao Jiang , Weimin Tan , Tian Tan , Bo Yan , Liquan Shen

Embodied navigation agents built upon large reasoning models (LRMs) can handle complex, multimodal environmental input and perform grounded reasoning per step to improve sequential decision-making for long-horizon tasks. However, a critical…

人工智能 · 计算机科学 2026-04-10 He Zhao , Yijun Yang , Zichuan Lin , Deheng Ye , Chunyan Miao

Large multimodal models (LMMs) show strong visual-linguistic reasoning but their capacity for spatial decision-making and action remains unclear. In this work, we investigate whether LMMs can achieve embodied spatial action like human…

Open-world promptable 3D semantic segmentation remains brittle as semantics are inferred in the input sensor coordinates. Yet, humans, in contrast, interpret parts via functional roles in a canonical space -- wings extend laterally, handles…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Li Jin , Weikai Chen , Yujie Wang , Yingda Yin , Zeyu Hu , Runze Zhang , Keyang Luo , Shengju Qian , Xin Wang , Xueying Qin

Vision language navigation is the task that requires an agent to navigate through a 3D environment based on natural language instructions. One key challenge in this task is to ground instructions with the current visual information that the…

计算与语言 · 计算机科学 2021-04-21 Jialu Li , Hao Tan , Mohit Bansal

Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve…

机器人学 · 计算机科学 2024-10-15 Xinxin Zhao , Wenzhe Cai , Likun Tang , Teng Wang

Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging. Recent research on enhancing language-guided navigation…

人工智能 · 计算机科学 2026-02-10 Changxin Huang , Lv Tang , Zhaohuan Zhan , Lisha Yu , Runhao Zeng , Zun Liu , Zhengjie Wang , Jianqiang Li

Audio-Visual Embodied Navigation aims to enable agents to autonomously navigate to sound sources in unknown 3D environments using auditory cues. While current AVN methods excel on in-distribution sound sources, they exhibit poor…

声音 · 计算机科学 2025-10-15 Yi Wang , Yinfeng Yu , Fuchun Sun , Liejun Wang , Wendong Zheng

With the emergence of varied visual navigation tasks (e.g, image-/object-/audio-goal and vision-language navigation) that specify the target in different ways, the community has made appealing advances in training specialized agents capable…

计算机视觉与模式识别 · 计算机科学 2022-11-01 Hanqing Wang , Wei Liang , Luc Van Gool , Wenguan Wang

Embodied Question Answering (EQA) is a recently proposed task, where an agent is placed in a rich 3D environment and must act based solely on its egocentric input to answer a given question. The desired outcome is that the agent learns to…

计算机视觉与模式识别 · 计算机科学 2019-08-15 Cătălina Cangea , Eugene Belilovsky , Pietro Liò , Aaron Courville

Image-goal navigation is a challenging task, as it requires the agent to navigate to a target indicated by an image in a previously unseen scene. Current methods introduce diverse memory mechanisms which save navigation history to solve…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Hongxin Li , Xu Yang , Yuran Yang , Shuqi Mei , Zhaoxiang Zhang

Understanding the geometric and semantic structure of environments is essential for embodied navigation and reasoning. Existing semantic mapping methods trade off between explicit geometry and multi-scale semantics, and lack a native…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Sixian Zhang , Yiyao Wang , Xinhang Song , Keming Zhang , Zijian Xu , Shuqiang Jiang

Camouflaged object detection (COD) aims to identify the objects that conceal themselves in natural scenes. Accurate COD suffers from a number of challenges associated with low boundary contrast and the large variation of object appearances,…

计算机视觉与模式识别 · 计算机科学 2022-07-28 Geng Chen , Si-Jie Liu , Yu-Jia Sun , Ge-Peng Ji , Ya-Feng Wu , Tao Zhou

Moving around in the world is naturally a multisensory experience, but today's embodied agents are deaf---restricted to solely their visual perception of the environment. We introduce audio-visual navigation for complex, acoustically and…

Cooperative perception, offering a wider field of view than standalone perception, is becoming increasingly crucial in autonomous driving. This perception is enabled through vehicle-to-vehicle (V2V) communication, allowing connected…

信息论 · 计算机科学 2024-09-17 Yucheng Sheng , Le Liang , Hao Ye , Shi Jin , Geoffrey Ye Li