中文
相关论文

相关论文: SLARM: Streaming and Language-Aligned Reconstructi…

200 篇论文

We propose SemGauss-SLAM, a dense semantic SLAM system utilizing 3D Gaussian representation, that enables accurate 3D semantic mapping, robust camera tracking, and high-quality rendering simultaneously. In this system, we incorporate…

机器人学 · 计算机科学 2025-06-25 Siting Zhu , Renjie Qin , Guangming Wang , Jiuming Liu , Hesheng Wang

In recent years, video semantic segmentation has made great progress with advanced deep neural networks. However, there still exist two main challenges \ie, information inconsistency and computation cost. To deal with the two difficulties,…

计算机视觉与模式识别 · 计算机科学 2023-04-19 Jinming Su , Ruihong Yin , Shuaibin Zhang , Junfeng Luo

We propose Hier-SLAM, a semantic 3D Gaussian Splatting SLAM method featuring a novel hierarchical categorical representation, which enables accurate global 3D semantic mapping, scaling-up capability, and explicit semantic label prediction…

机器人学 · 计算机科学 2025-03-11 Boying Li , Zhixi Cai , Yuan-Fang Li , Ian Reid , Hamid Rezatofighi

Real-time reconstruction of dynamic 3D scenes from uncalibrated video streams demands robust online methods that recover scene dynamics from sparse observations under strict latency and memory constraints. Yet most dynamic reconstruction…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Zike Wu , Qi Yan , Xuanyu Yi , Lele Wang , Renjie Liao

Holistic 3D scene understanding, which jointly models geometry, appearance, and semantics, is crucial for applications like augmented reality and robotic interaction. Existing feed-forward 3D scene understanding methods (e.g., LSM) are…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Qijing Li , Jingxiang Sun , Liang An , Zhaoqi Su , Hongwen Zhang , Yebin Liu

We present STORM (Search-Guided Generative World Models), a novel framework for spatio-temporal reasoning in robotic manipulation that unifies diffusion-based action generation, conditional video prediction, and search-based planning.…

机器人学 · 计算机科学 2025-12-23 Wenjun Lin , Jensen Zhang , Kaitong Cai , Keze Wang

Context-aware STR methods typically use internal autoregressive (AR) language models (LM). Inherent limitations of AR models motivated two-stage methods which employ an external LM. The conditional independence of the external LM on the…

计算机视觉与模式识别 · 计算机科学 2022-07-15 Darwin Bautista , Rowel Atienza

Composed image retrieval (CIR) is a vision language task that retrieves a target image using a reference image and modification text, enabling intuitive specification of desired changes. While effectively fusing visual and textual…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Jeong-Woo Park , Young-Eun Kim , Seong-Whan Lee

With the success of pre-trained visual-language (VL) models such as CLIP in visual representation tasks, transferring pre-trained models to downstream tasks has become a crucial paradigm. Recently, the prompt tuning paradigm, which draws…

计算机视觉与模式识别 · 计算机科学 2023-12-14 Jingsheng Gao , Jiacheng Ruan , Suncheng Xiang , Zefang Yu , Ke Ji , Mingye Xie , Ting Liu , Yuzhuo Fu

We propose a novel, vision-only object-level SLAM framework for automotive applications representing 3D shapes by implicit signed distance functions. Our key innovation consists of augmenting the standard neural representation by a…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Li Cui , Yang Ding , Richard Hartley , Zirui Xie , Laurent Kneip , Zhenghua Yu

Emergent communication enables partially observant Autonomous Mobile Robots (AMRs) to coordinate effectively in decentralized multi-agent reinforcement learning (MARL) settings. However, existing approaches often struggle with unstable…

机器人学 · 计算机科学 2026-05-28 Mahmoud Abouelyazid , Eman Hammad

CLIP has shown impressive results in aligning images and texts at scale. However, its ability to capture detailed visual features remains limited because CLIP matches images and texts at a global level. To address this issue, we propose…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Rui Xiao , Sanghwan Kim , Mariana-Iuliana Georgescu , Zeynep Akata , Stephan Alaniz

In the field of SLAM (Simultaneous Localization And Mapping) for robot navigation, mapping the environment is an important task. In this regard the Lidar sensor can produce near accurate 3D map of the environment in the format of point…

计算机视觉与模式识别 · 计算机科学 2020-09-15 Aritra Mukherjee , Sourya Dipta Das , Jasorsi Ghosh , Ananda S. Chowdhury , Sanjoy Kumar Saha

3D human reconstruction and animation are long-standing topics in computer graphics and vision. However, existing methods typically rely on sophisticated dense-view capture and/or time-consuming per-subject optimization procedures. To…

图形学 · 计算机科学 2025-06-04 Zhiyuan Yu , Zhe Li , Hujun Bao , Can Yang , Xiaowei Zhou

Attention mechanisms and non-local mean operations in general are key ingredients in many state-of-the-art deep learning techniques. In particular, the Transformer model based on multi-head self-attention has recently achieved great success…

机器学习 · 计算机科学 2019-05-27 Dan A. Calian , Peter Roelants , Jacques Cali , Ben Carr , Krishna Dubba , John E. Reid , Dell Zhang

Neural implicit representations have recently shown encouraging results in various domains, including promising progress in simultaneous localization and mapping (SLAM). Nevertheless, existing methods produce over-smoothed scene…

计算机视觉与模式识别 · 计算机科学 2022-04-22 Zihan Zhu , Songyou Peng , Viktor Larsson , Weiwei Xu , Hujun Bao , Zhaopeng Cui , Martin R. Oswald , Marc Pollefeys

Traditional approaches to stereo visual SLAM rely on point features to estimate the camera trajectory and build a map of the environment. In low-textured environments, though, it is often difficult to find a sufficient number of reliable…

计算机视觉与模式识别 · 计算机科学 2018-04-10 Ruben Gomez-Ojeda , David Zuñiga-Noël , Francisco-Angel Moreno , Davide Scaramuzza , Javier Gonzalez-Jimenez

Simultaneous Localization and Mapping (SLAM) is a critical task in robotics, enabling systems to autonomously navigate and understand complex environments. Current SLAM approaches predominantly rely on geometric cues for mapping and…

机器人学 · 计算机科学 2025-03-28 Yongxu Wang , Xu Cao , Weiyun Yi , Zhaoxin Fan

While Multimodal Large Language Models (MLLMs) excel in semantic tasks, they frequently lack the "spatial sense" essential for sophisticated geometric reasoning. Current models typically suffer from exorbitant modality-alignment costs and…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Yi Zhang , Youya Xia , Yong Wang , Meng Song , Xin Wu , Wenjun Wan , Bingbing Liu , AiXue Ye , Hongbo Zhang , Feng Wen

The choice of scene representation is crucial in both the shape inference algorithms it requires and the smart applications it enables. We present efficient and optimisable multi-class learned object descriptors together with a novel…

计算机视觉与模式识别 · 计算机科学 2020-10-13 Edgar Sucar , Kentaro Wada , Andrew Davison