English
Related papers

Related papers: When Pedestrian Detection Meets Multi-Modal Learni…

200 papers

Jointly processing information from multiple sensors is crucial to achieving accurate and robust perception for reliable autonomous driving systems. However, current 3D perception research follows a modality-specific paradigm, leading to…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Haiyang Wang , Hao Tang , Shaoshuai Shi , Aoxue Li , Zhenguo Li , Bernt Schiele , Liwei Wang

We propose a deep neural network fusion architecture for fast and robust pedestrian detection. The proposed network fusion architecture allows for parallel processing of multiple networks for speed. A single shot deep convolutional network…

Computer Vision and Pattern Recognition · Computer Science 2017-05-30 Xianzhi Du , Mostafa El-Khamy , Jungwon Lee , Larry S. Davis

Crowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited information of RGB images and cannot well discover…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Lingbo Liu , Jiaqi Chen , Hefeng Wu , Guanbin Li , Chenglong Li , Liang Lin

Recently, many multi-modal trackers prioritize RGB as the dominant modality, treating other modalities as auxiliary, and fine-tuning separately various multi-modal tasks. This imbalance in modality dependence limits the ability of methods…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Xiantao Hu , Bineng Zhong , Qihua Liang , Zhiyi Mo , Liangtao Shi , Ying Tai , Jian Yang

In remote sensing, each sensor can provide complementary or reinforcing information. It is valuable to fuse outputs from multiple sensors to boost overall performance. Previous supervised fusion methods often require accurate labels for…

Computer Vision and Pattern Recognition · Computer Science 2019-11-22 Xiaoxiao Du , Alina Zare

Deep Learning has implemented a wide range of applications and has become increasingly popular in recent years. The goal of multimodal deep learning (MMDL) is to create models that can process and link information using various modalities.…

Machine Learning · Computer Science 2022-02-21 Jabeen Summaira , Xi Li , Amin Muhammad Shoib , Jabbar Abdul

In this study, we introduce AV-PedAware, a self-supervised audio-visual fusion system designed to improve dynamic pedestrian awareness for robotics applications. Pedestrian awareness is a critical requirement in many robotics applications.…

Robotics · Computer Science 2025-04-07 Yizhuo Yang , Shenghai Yuan , Muqing Cao , Jianfei Yang , Lihua Xie

Perceiving pedestrians in highly crowded urban environments is a difficult long-tail problem for learning-based autonomous perception. Speeding up 3D ground truth generation for such challenging scenes is performance-critical yet very…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Shichao Li , Peiliang Li , Qing Lian , Peng Yun , Xiaozhi Chen

With the proliferation of low altitude unmanned aerial vehicles (UAVs), visual multi-object tracking is becoming a critical security technology, demanding significant robustness even in complex environmental conditions. However, tracking…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Tianyang Xu , Jinjie Gu , Xuefeng Zhu , XiaoJun Wu , Josef Kittler

Multimodal VAEs seek to model the joint distribution over heterogeneous data (e.g.\ vision, language), whilst also capturing a shared representation across such modalities. Prior work has typically combined information from the modalities…

Machine Learning · Computer Science 2022-12-19 Tom Joy , Yuge Shi , Philip H. S. Torr , Tom Rainforth , Sebastian M. Schmon , N. Siddharth

Multi-modal collaborative perception calls for great attention to enhancing the safety of autonomous driving. However, current multi-modal approaches remain a ``local fusion to communication'' sequence, which fuses multi-modal data locally…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Kang Yang , Peng Wang , Lantao Li , Tianci Bu , Chen Sun , Deying Li , Yongcai Wang

During the process of driving, humans usually rely on multiple senses to gather information and make decisions. Analogously, in order to achieve embodied intelligence in autonomous driving, it is essential to integrate multidimensional…

Understanding and predicting the intention of pedestrians is essential to enable autonomous vehicles and mobile robots to navigate crowds. This problem becomes increasingly complex when we consider the uncertainty and multimodality of…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Stuart Eiffert , Kunming Li , Mao Shan , Stewart Worrall , Salah Sukkarieh , Eduardo Nebot

Pedestrian detection in crowded scenes is a challenging problem, because occlusion happens frequently among different pedestrians. In this paper, we propose an effective and efficient detection network to hunt pedestrians in crowd scenes.…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Cheng Chi , Shifeng Zhang , Junliang Xing , Zhen Lei , Stan Z. Li , Xudong Zou

Many keypoint detection and description methods have been proposed for image matching or registration. While these methods demonstrate promising performance for single-modality image matching, they often struggle with multimodal data…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Yepeng Liu , Zhichao Sun , Baosheng Yu , Yitian Zhao , Bo Du , Yongchao Xu , Jun Cheng

Current Pedestrian Attribute Recognition (PAR) algorithms typically focus on mapping visual features to semantic labels or attempt to enhance learning by fusing visual and attribute information. However, these methods fail to fully exploit…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Xiao Wang , Shujuan Wu , Xiaoxia Cheng , Changwei Bi , Jin Tang , Bin Luo

Multispectral images consisting of aligned visual-optical (VIS) and thermal infrared (IR) image pairs are well-suited for practical applications like autonomous driving or visual surveillance. Such data can be used to increase the…

Computer Vision and Pattern Recognition · Computer Science 2020-08-21 Alexander Wolpert , Michael Teutsch , M. Saquib Sarfraz , Rainer Stiefelhagen

Generalizable person re-identification (Re-ID) is a very hot research topic in machine learning and computer vision, which plays a significant role in realistic scenarios due to its various applications in public security and video…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Suncheng Xiang , Jingsheng Gao , Mengyuan Guan , Jiacheng Ruan , Chengfeng Zhou , Ting Liu , Dahong Qian , Yuzhuo Fu

Multimodal fusion frameworks for Human Action Recognition (HAR) using depth and inertial sensor data have been proposed over the years. In most of the existing works, fusion is performed at a single level (feature level or decision level),…

Machine Learning · Computer Science 2019-10-28 Zeeshan Ahmad , Naimul Khan

Despite the growing popularity of Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algorithmic progress or are artifacts of inconsistent evaluation…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Hao Dong , Hongzhao Li , Shupan Li , Muhammad Haris Khan , Eleni Chatzi , Olga Fink