English
Related papers

Related papers: Multi-Scale Gaussian-Language Map for Zero-shot Em…

200 papers

Multimodal Small-to-Medium sized Language Models (MSLMs) have demonstrated strong capabilities in integrating visual and textual information but still face significant limitations in visual comprehension and mathematical reasoning,…

Machine Learning · Computer Science 2026-01-27 Ashutosh Bajpai , Akshat Bhandari , Akshay Nambi , Tanmoy Chakraborty

Cooperative visual semantic navigation is a foundational capability for aerial robot teams operating in unknown environments. However, achieving robust open-vocabulary object-goal navigation remains challenging due to the computational…

Robotics · Computer Science 2026-03-17 MoniJesu Wonders James , Amir Atef Habel , Aleksey Fedoseev , Dzmitry Tsetserokou

Embodied agents equipped with GPT as their brains have exhibited extraordinary decision-making and generalization abilities across various tasks. However, existing zero-shot agents for vision-and-language navigation (VLN) only prompt GPT-4…

Artificial Intelligence · Computer Science 2024-06-21 Jiaqi Chen , Bingqian Lin , Ran Xu , Zhenhua Chai , Xiaodan Liang , Kwan-Yee K. Wong

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, which requires multi-round dialogue spatial reasoning and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Xunyi Zhao , Gengze Zhou , Qi Wu

We propose NEDS-SLAM, a dense semantic SLAM system based on 3D Gaussian representation, that enables robust 3D semantic mapping, accurate camera tracking, and high-quality rendering in real-time. In the system, we propose a Spatially…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Yiming Ji , Yang Liu , Guanghu Xie , Boyu Ma , Zongwu Xie

Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often struggle to accurately interpret depth information inherent in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Hao Yang , Hongbo Zhang , Yanyan Zhao , Bing Qin

We present an active mapping system that plans for both long-horizon exploration goals and short-term actions using a 3D Gaussian Splatting (3DGS) representation. Existing methods either do not take advantage of recent developments in…

Robotics · Computer Science 2025-09-08 Wen Jiang , Boshu Lei , Katrina Ashton , Kostas Daniilidis

We present SGS-SLAM, the first semantic visual SLAM system based on Gaussian Splatting. It incorporates appearance, geometry, and semantic features through multi-channel optimization, addressing the oversmoothing limitations of neural…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Mingrui Li , Shuhong Liu , Heng Zhou , Guohao Zhu , Na Cheng , Tianchen Deng , Hongyu Wang

Visual navigation with an image as goal is a fundamental and challenging problem. Conventional methods either rely on end-to-end RL learning or modular-based policy with topological graph or BEV map as memory, which cannot fully model the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Wenxuan Guo , Xiuwei Xu , Hang Yin , Ziwei Wang , Jianjiang Feng , Jie Zhou , Jiwen Lu

We propose a framework for active mapping and exploration that leverages Gaussian splatting for constructing dense maps. Further, we develop a GPU-accelerated motion planning algorithm that can exploit the Gaussian map for real-time…

Robotics · Computer Science 2025-10-07 Yuezhan Tao , Dexter Ong , Varun Murali , Igor Spasojevic , Pratik Chaudhari , Vijay Kumar

Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly,…

Large language models (LLMs) have shown promising results in learning and contextualizing information from different forms of data. Recent advancements in foundational models, particularly those employing self-attention mechanisms, have…

Computation and Language · Computer Science 2024-07-17 Devashish Vikas Gupta , Azeez Syed Ali Ishaqui , Divya Kiran Kadiyala

Simultaneous Localization and Mapping (SLAM) techniques play a key role towards long-term autonomy of mobile robots due to the ability to correct localization errors and produce consistent maps of an environment over time. Contrarily to…

While Multimodal Large Language Models (MLLMs) have achieved remarkable success in 2D visual understanding, their ability to reason about 3D space remains limited. To address this gap, we introduce geometrically referenced 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Jiangye Yuan , Gowri Kumar , Baoyuan Wang

Multimodal Large Language Models (MLLMs) have achieved remarkable progress but continue to struggle with geometric reasoning, primarily due to the perception bottleneck regarding fine-grained visual elements. While formal languages have…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Peijie Wang , Ming-Liang Zhang , Jun Cao , Chao Deng , Dekang Ran , Hongda Sun , Pi Bu , Xuan Zhang , Yingyao Wang , Jun Song , Bo Zheng , Fei Yin , Cheng-Lin Liu

The creation of a metric-semantic map, which encodes human-prior knowledge, represents a high-level abstraction of environments. However, constructing such a map poses challenges related to the fusion of multi-modal sensor data, the…

Robotics · Computer Science 2024-12-03 Jianhao Jiao , Ruoyu Geng , Yuanhang Li , Ren Xin , Bowen Yang , Jin Wu , Lujia Wang , Ming Liu , Rui Fan , Dimitrios Kanoulas

Localization and mapping are critical tasks for various applications such as autonomous vehicles and robotics. The challenges posed by outdoor environments present particular complexities due to their unbounded characteristics. In this…

Robotics · Computer Science 2024-04-08 Chenyang Wu , Yifan Duan , Xinran Zhang , Yu Sheng , Jianmin Ji , Yanyong Zhang

Intelligent embodied agents (e.g. robots) need to perform complex semantic tasks in unfamiliar environments. Among many skills that the agents need to possess, building and maintaining a semantic map of the environment is most crucial in…

Robotics · Computer Science 2025-08-13 Sonia Raychaudhuri , Angel X. Chang

While interacting in the world is a multi-sensory experience, many robots continue to predominantly rely on visual perception to map and navigate in their environments. In this work, we propose Audio-Visual-Language Maps (AVLMaps), a…

Robotics · Computer Science 2023-03-28 Chenguang Huang , Oier Mees , Andy Zeng , Wolfram Burgard

This paper proposes Neural-MMGS, a novel neural 3DGS framework for multimodal large-scale scene reconstruction that fuses multiple sensing modalities in a per-gaussian compact, learnable embedding. While recent works focusing on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Sitian Shen , Georgi Pramatarov , Yifu Tao , Daniele De Martini