English
Related papers

Related papers: Forging Spatial Intelligence: A Roadmap of Multi-M…

200 papers

Multi-modal learning is a fast growing area in artificial intelligence. It tries to help machines understand complex things by combining information from different sources, like images, text, and audio. By using the strengths of each…

Machine Learning · Computer Science 2025-12-22 Qihang Jin , Enze Ge , Yuhang Xie , Hongying Luo , Junhao Song , Ziqian Bi , Chia Xin Liang , Jibin Guan , Joe Yeong , Xinyuan Song , Junfeng Hao

Autonomous driving sensors generate an enormous amount of data. In this paper, we explore learned multimodal compression for autonomous driving, specifically targeted at 3D object detection. We focus on camera and LiDAR modalities and…

Image and Video Processing · Electrical Eng. & Systems 2024-08-16 Hadi Hadizadeh , Ivan V. Bajić

Recent advancements in foundation models, typically trained with self-supervised learning on large-scale and diverse datasets, have shown great potential in medical image analysis. However, due to the significant spatial heterogeneity of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Lingxiao Luo , Xuanzhong Chen , Bingda Tang , Xinsheng Chen , Rong Han , Chengpeng Hu , Yujiang Li , Ting Chen

Building a multi-modality multi-task neural network toward accurate and robust performance is a de-facto standard in perception task of autonomous driving. However, leveraging such data from multiple sensors to jointly optimize the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Tengju Ye , Wei Jing , Chunyong Hu , Shikun Huang , Lingping Gao , Fangzhen Li , Jingke Wang , Ke Guo , Wencong Xiao , Weibo Mao , Hang Zheng , Kun Li , Junbo Chen , Kaicheng Yu

Deep learning based localization and mapping has recently attracted significant attention. Instead of creating hand-designed algorithms through exploitation of physical models or geometric theories, deep learning based solutions provide an…

Computer Vision and Pattern Recognition · Computer Science 2020-07-01 Changhao Chen , Bing Wang , Chris Xiaoxuan Lu , Niki Trigoni , Andrew Markham

Grounding language to a navigating agent's observations can leverage pretrained multimodal foundation models to match perceptions to object or event descriptions. However, previous approaches remain disconnected from environment mapping,…

Robotics · Computer Science 2025-06-10 Chenguang Huang , Oier Mees , Andy Zeng , Wolfram Burgard

The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models remains opaque. We provide empirical clarity through…

Mapping and navigation have gone hand-in-hand since long before robots existed. Maps are a key form of communication, allowing someone who has never been somewhere to nonetheless navigate that area successfully. In the context of…

Robotics · Computer Science 2024-07-16 Ian D. Miller , Fernando Cladera , Trey Smith , Camillo Jose Taylor , Vijay Kumar

The last decade witnessed increasingly rapid progress in self-driving vehicle technology, mainly backed up by advances in the area of deep learning and artificial intelligence. The objective of this paper is to survey the current…

Machine Learning · Computer Science 2020-03-26 Sorin Grigorescu , Bogdan Trasnea , Tiberiu Cocias , Gigel Macesanu

Federated learning involves training statistical models over edge devices such as mobile phones such that the training data is kept local. Federated Learning (FL) can serve as an ideal candidate for training spatial temporal models that…

Machine Learning · Computer Science 2024-02-09 Yacine Belal , Sonia Ben Mokhtar , Hamed Haddadi , Jaron Wang , Afra Mashhadi

Foundation models have transformed natural language processing and computer vision, and their impact is now reshaping remote sensing image analysis. With powerful generalization and transfer learning capabilities, they align naturally with…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Liling Yang , Ning Chen , Jun Yue , Yidan Liu , Jiayi Ma , Pedram Ghamisi , Antonio Plaza , Leyuan Fang

Robust semantic perception for autonomous vehicles relies on effectively combining multiple sensors with complementary strengths and weaknesses. State-of-the-art sensor fusion approaches to semantic perception often treat sensor data…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Tim Broedermannn , Christos Sakaridis , Luigi Piccinelli , Wim Abbeloos , Luc Van Gool

Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We adopt a task-oriented perspective to systematically review the applications and advancements of multimodal fusion…

Rapid advancements in foundation models, including Large Language Models, Vision-Language Models, Multimodal Large Language Models, and Vision-Language-Action Models, have opened new avenues for embodied AI in mobile service robotics. By…

Robotics · Computer Science 2026-03-11 Matthew Lisondra , Beno Benhabib , Goldie Nejat

Interpreting natural-language commands to localize target objects is critical for autonomous driving (AD). Existing visual grounding (VG) methods for autonomous vehicles (AVs) typically struggle with ambiguous, context-dependent…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Haicheng Liao , Huanming Shen , Bonan Wang , Yongkang Li , Yihong Tang , Chengyue Wang , Dingyi Zhuang , Kehua Chen , Hai Yang , Chengzhong Xu , Zhenning Li

As an effective approach to understanding the human-centric physical world, Wearable Artificial Intelligence (AI), which leverages multimodal wearable sensors to understand human physiology and behavior, has attracted increasing attention…

Signal Processing · Electrical Eng. & Systems 2026-04-14 Yize Cai , Baoshen Guo , Guobin Shen , Zhiqing Hong

Autonomous driving is a popular research area within the computer vision research community. Since autonomous vehicles are highly safety-critical, ensuring robustness is essential for real-world deployment. While several public multimodal…

Spatial grounding, the process of associating natural language expressions with corresponding image regions, has rapidly advanced due to the introduction of transformer-based models, significantly enhancing multimodal representation and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Ijazul Haq , Muhammad Saqib , Yingjie Zhang

This paper presents a novel layered framework that integrates visual foundation models to improve robot manipulation tasks and motion planning. The framework consists of five layers: Perception, Cognition, Planning, Execution, and Learning.…

Robotics · Computer Science 2023-09-21 Chen Yang , Peng Zhou , Jiaming Qi

Geometric navigation is nowadays a well-established field of robotics and the research focus is shifting towards higher-level scene understanding, such as Semantic Mapping. When a robot needs to interact with its environment, it must be…

Robotics · Computer Science 2023-11-23 Federico Rollo , Gennaro Raiola , Andrea Zunino , Nikolaos Tsagarakis , Arash Ajoudani
‹ Prev 1 3 4 5 6 7 10 Next ›