English
Related papers

Related papers: MEM: Multi-Modal Elevation Mapping for Robotics an…

200 papers

Developing effective multimodal data fusion strategies has become increasingly essential for improving the predictive power of statistical machine learning methods across a wide range of applications, from autonomous driving to medical…

Machine Learning · Computer Science 2025-07-29 Ziyi Liang , Annie Qu , Babak Shahbaba

Heterogeneous multi-robot systems feature significant adaptability for complex environments. However, effective collaboration that fully exploits the robots' potential remains a core challenge. This paper proposes a decentralized…

Robotics · Computer Science 2026-05-12 Yuxiang Li , Kun Chen , Jiancheng Wang , Shihao Fang , Haoyao Chen , Yunhui Liu

Scaling large multimodal models (LMMs) to 3D understanding poses unique challenges: point cloud data is sparse and irregular, existing models rely on fragmented architectures with modality-specific encoders, and training pipelines often…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Yongyuan Liang , Xiyao Wang , Yuanchen Ju , Jianwei Yang , Furong Huang

This paper presents a strategy to guide a mobile ground robot equipped with a camera or depth sensor, in order to autonomously map the visible part of a bounded three-dimensional structure. We describe motion planning algorithms that…

Robotics · Computer Science 2017-11-15 Manikandasriram Srinivasan Ramanagopal , André Phu-Van Nguyen , Jerome Le Ny

Recent advancements in perception for autonomous driving are driven by deep learning. In order to achieve robust and accurate scene understanding, autonomous vehicles are usually equipped with different sensors (e.g. cameras, LiDARs,…

Robotic manipulation and navigation are fundamental capabilities of embodied intelligence, enabling effective robot interactions with the physical world. Achieving these capabilities requires a cohesive understanding of the environment,…

Robotics · Computer Science 2025-11-18 Xiaoshuai Hao , Yingbo Tang , Lingfeng Zhang , Yanbiao Ma , Yunfeng Diao , Ziyu Jia , Wenbo Ding , Hangjun Ye , Long Chen

Learning contextual and spatial environmental representations enhances autonomous vehicle's hazard anticipation and decision-making in complex scenarios. Recent perception systems enhance spatial understanding with sensor fusion but often…

Robotics · Computer Science 2024-01-18 Shoaib Azam , Farzeen Munir , Ville Kyrki , Moongu Jeon , Witold Pedrycz

Multimodal summarization integrating information from diverse data modalities presents a promising solution to aid the understanding of information within various processes. However, the application and advantages of multimodal…

Software Engineering · Computer Science 2025-03-07 Nenad Petrovic , Yurui Zhang , Moaad Maaroufi , Kuo-Yi Chao , Lukasz Mazur , Fengjunjie Pan , Vahid Zolfaghari , Alois Knoll

Classification and identification of the materials lying over or beneath the Earth's surface have long been a fundamental but challenging research topic in geoscience and remote sensing (RS) and have garnered a growing concern owing to the…

Computer Vision and Pattern Recognition · Computer Science 2020-08-13 Danfeng Hong , Lianru Gao , Naoto Yokoya , Jing Yao , Jocelyn Chanussot , Qian Du , Bing Zhang

Safe and efficient robot operation in complex human environments can benefit from good models of site-specific motion patterns. Maps of Dynamics (MoDs) provide such models by encoding statistical motion patterns in a map, but existing…

This paper presents a Multi-Elevation Semantic Segmentation Image (MESSI) dataset comprising 2525 images taken by a drone flying over dense urban environments. MESSI is unique in two main features. First, it contains images from various…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Barak Pinkovich , Boaz Matalon , Ehud Rivlin , Hector Rotstein

The advent of generalist Large Language Models (LLMs) and Large Vision Models (VLMs) have streamlined the construction of semantically enriched maps that can enable robots to ground high-level reasoning and planning into their…

Robotics · Computer Science 2024-11-06 Emilio Olivastri , Jonathan Francis , Alberto Pretto , Niko Sünderhauf , Krishan Rana

Beamforming (BF) is essential for enhancing system capacity in fifth generation (5G) and beyond wireless networks, yet exhaustive beam training in ultra-massive multiple-input multiple-output (MIMO) systems incurs substantial overhead. To…

Signal Processing · Electrical Eng. & Systems 2026-02-11 Yanliang Jin , Yunfan Li , Jiang Jun , Yuan Gao , Shengli Liu , Jianbo Du , Zhaohui Yang , Shugong Xu

Boundary information plays a significant role in 2D image segmentation, while usually being ignored in 3D point cloud segmentation where ambiguous features might be generated in feature extraction, leading to misclassification in the…

Computer Vision and Pattern Recognition · Computer Science 2021-01-08 Jingyu Gong , Jiachen Xu , Xin Tan , Jie Zhou , Yanyun Qu , Yuan Xie , Lizhuang Ma

The goal of multimodal alignment is to learn a single latent space that is shared between multimodal inputs. The most powerful models in this space have been trained using massive datasets of paired inputs and large-scale computational…

The rapid increase in multimedia data has spurred advancements in Multimodal Summarization with Multimodal Output (MSMO), which aims to produce a multimodal summary that integrates both text and relevant images. The inherent heterogeneity…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Yanghai Zhang , Ye Liu , Shiwei Wu , Kai Zhang , Xukai Liu , Qi Liu , Enhong Chen

Deploying depth estimation networks in the real world requires high-level robustness against various adverse conditions to ensure safe and reliable autonomy. For this purpose, many autonomous vehicles employ multi-modal sensor systems,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Ukcheol Shin , Kyunghyun Lee , Jean Oh

Precise perception of articulated objects is vital for empowering service robots. Recent studies mainly focus on point cloud, a single-modal approach, often neglecting vital texture and lighting details and assuming ideal conditions like…

Robotics · Computer Science 2024-07-02 Hongliang Zeng , Ping Zhang , Chengjiong Wu , Jiahua Wang , Tingyu Ye , Fang Li

Multi-modal image fusion (MMIF) maps useful information from various modalities into the same representation space, thereby producing an informative fused image. However, the existing fusion algorithms tend to symmetrically fuse the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Jingxue Huang , Xilai Li , Tianshu Tan , Xiaosong Li , Tao Ye

The rapid growth in terms of the availability of transportation data provides great potential for the introduction of emerging data-driven methodologies into transportation-related research and development efforts. However, advanced…

Physics and Society · Physics 2024-06-25 Zilin Bian , Dachuan Zuo , Jingqin Gao , Kaan Ozbay , Matthew D. Maggio