English
Related papers

Related papers: MarS3D: A Plug-and-Play Motion-Aware Model for Sem…

200 papers

Localization is an essential task for mobile autonomous robotic systems that want to use pre-existing maps or create new ones in the context of SLAM. Today, many robotic platforms are equipped with high-accuracy 3D LiDAR sensors, which…

Point cloud stands as the most widely adopted format for representing 3D shapes and scenes due to its simplicity and geometric fidelity. However, its inherent unordered and irregular nature, exacerbated by sensor noise and occlusions,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Minhas Kamal , Hiranya Garbha Kumar , Balakrishnan Prabhakaran

2D images and 3D point clouds are foundational data types for multimedia applications, including real-time video analysis, augmented reality (AR), and 3D scene understanding. Class-incremental semantic segmentation (CSS) requires…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Jiaxu Li , Rui Li , Jianyu Qi , Songning Lai , Linpu Lv , Kejia Fan , Jianheng Tang , Yutao Yue , Dongzhan Zhou , Yuanhuai Liu , Huiping Zhuang

3D scene understanding is a critical yet challenging task in autonomous driving due to the irregularity and sparsity of LiDAR data, as well as the computational demands of processing large-scale point clouds. Recent methods leverage…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Bin Yang , Alexandru Paul Condurache

Vision-based 3D semantic occupancy prediction is a critical task in 3D vision that integrates volumetric 3D reconstruction with semantic understanding. Existing methods, however, often rely on modular pipelines. These modules are typically…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Dubing Chen , Huan Zheng , Yucheng Zhou , Xianfei Li , Wenlong Liao , Tao He , Pai Peng , Jianbing Shen

3D point cloud semantic segmentation is a challenging topic in the computer vision field. Most of the existing methods in literature require a large amount of fully labeled training data, but it is extremely time-consuming to obtain these…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Shuang Deng , Qiulei Dong , Bo Liu , Zhanyi Hu

Image segmentation plays a pivotal role in several medical-imaging applications by assisting the segmentation of the regions of interest. Deep learning-based approaches have been widely adopted for semantic segmentation of medical data. In…

Image and Video Processing · Electrical Eng. & Systems 2021-01-20 Abhishek Shivdeo , Rohit Lokwani , Viraj Kulkarni , Amit Kharat , Aniruddha Pant

Open-vocabulary semantic mapping enables robots to spatially ground previously unseen concepts without requiring predefined class sets. Current training-free methods commonly rely on multi-view fusion of semantic embeddings into a 3D map,…

Accurate segmentation of 3D medical images is critical for clinical applications like disease assessment and treatment planning. While the Segment Anything Model 2 (SAM2) has shown remarkable success in video object segmentation by…

Image and Video Processing · Electrical Eng. & Systems 2025-10-13 Yeqing Yang , Le Xu , Lixia Tian

Point clouds are a set of data points in space to represent the 3D geometry of objects. A fundamental step in the processing is to identify a subset of points to represent the shape. While traditional sampling methods often ignore to…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Pierre Onghena , Santiago Velasco-Forero , Beatriz Marcotegui

Open-vocabulary 3D scene understanding presents a significant challenge in computer vision, with wide-ranging applications in embodied agents and augmented reality systems. Existing methods adopt neurel rendering methods as 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Jun Guo , Xiaojian Ma , Yue Fan , Huaping Liu , Qing Li

Point cloud segmentation is a fundamental task in 3D scene understanding. Its progress is constrained by the high cost and time required for dense 3D annotations, making labeled samples difficult to obtain. Beyond annotation scarcity,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Thenukan Pathmanathan , Kanchan Keisham , Thangarajah Akilan

For applications such as autonomous driving, self-localization/camera pose estimation and scene parsing are crucial technologies. In this paper, we propose a unified framework to tackle these two problems simultaneously. The uniqueness of…

Computer Vision and Pattern Recognition · Computer Science 2018-09-26 Peng Wang , Ruigang Yang , Binbin Cao , Wei Xu , Yuanqing Lin

Multi-modal 3D semantic segmentation is vital for applications such as autonomous driving and virtual reality (VR). To effectively deploy these models in real-world scenarios, it is essential to employ cross-domain adaptation techniques…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Mingyu Yang , Jitong Lu , Hun-Seok Kim

Robots are expected to operate autonomously in dynamic environments. Understanding the underlying dynamic characteristics of objects is a key enabler for achieving this goal. In this paper, we propose a method for pointwise semantic…

Robotics · Computer Science 2017-06-27 Ayush Dewan , Gabriel L. Oliveira , Wolfram Burgard

In the field of medical imaging, AI-assisted techniques such as object detection, segmentation, and classification are widely employed to alleviate the workload of physicians and doctors. However, single-task models are predominantly used,…

Image and Video Processing · Electrical Eng. & Systems 2025-11-18 Fan Li , Arun Iyengar , Lanyu Xu

Feature matching is a fundamental problem in computer vision with wide-ranging applications, including simultaneous localization and mapping (SLAM), image stitching, and 3D reconstruction. While recent advances in deep learning have…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ronald Nap , Andy Xiao

The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Pengcheng Zhao , Jinxing Zhou , Yang Zhao , Dan Guo , Yanxiang Chen

Understanding functionalities in 3D scenes involves interpreting natural language descriptions to locate functional interactive objects, such as handles and buttons, in a 3D environment. Functionality understanding is highly challenging, as…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Jaime Corsetti , Francesco Giuliari , Alice Fasoli , Davide Boscaini , Fabio Poiesi

Autonomous robotic systems and self driving cars rely on accurate perception of their surroundings as the safety of the passengers and pedestrians is the top priority. Semantic segmentation is one the essential components of environmental…

Computer Vision and Pattern Recognition · Computer Science 2021-02-10 Ran Cheng , Ryan Razani , Ehsan Taghavi , Enxu Li , Bingbing Liu