中文
相关论文

相关论文: Revisiting Modality Imbalance In Multimodal Pedest…

200 篇论文

Multi-modal ophthalmic image classification plays a key role in diagnosing eye diseases, as it integrates information from different sources to complement their respective performances. However, recent improvements have mainly focused on…

图像与视频处理 · 电气工程与系统科学 2024-05-29 Ke Zou , Tian Lin , Zongbo Han , Meng Wang , Xuedong Yuan , Haoyu Chen , Changqing Zhang , Xiaojing Shen , Huazhu Fu

We propose a novel method, Modality-based Redundancy Reduction Fusion (MRRF), for understanding and modulating the relative contribution of each modality in multimodal inference tasks. This is achieved by obtaining an $(M+1)$-way tensor to…

机器学习 · 计算机科学 2023-04-18 Elham J. Barezi , Peyman Momeni , Pascale Fung

Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the success of separate-encoder approaches like CLIP align…

计算与语言 · 计算机科学 2025-10-20 Qiyu Wu , Shuyang Cui , Satoshi Hayakawa , Wei-Yao Wang , Hiromi Wakaki , Yuki Mitsufuji

Combining complementary information from multiple modalities is intuitively appealing for improving the performance of learning-based approaches. However, it is challenging to fully leverage different modalities due to practical challenges…

机器学习 · 统计学 2018-05-31 Kuan Liu , Yanen Li , Ning Xu , Prem Natarajan

Visible-infrared person re-identification (VI-ReID) aims to retrieve images of the same pedestrian from different modalities, where the challenges lie in the significant modality discrepancy. To alleviate the modality gap, recent methods…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Zhihao Qian , Yutian Lin , Bo Du

Sensor fusion is crucial for a performant and robust Perception system in autonomous vehicles, but sensor staleness, where data from different sensors arrives with varying delays, poses significant challenges. Temporal misalignment between…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Meng Fan , Yifan Zuo , Patrick Blaes , Harley Montgomery , Subhasis Das

Standard multi-modal models assume the use of the same modalities in training and inference stages. However, in practice, the environment in which multi-modal models operate may not satisfy such assumption. As such, their performances…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Sangmin Woo , Sumin Lee , Yeonju Park , Muhammad Adi Nugroho , Changick Kim

The purpose of training neural networks is to achieve high generalization performance on unseen inputs. However, when trained on imbalanced datasets, a model's prediction tends to favor majority classes over minority classes, leading to…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Hiroaki Aizawa , Yuta Naito , Kohei Fukuda

Modality gap significantly restricts the effectiveness of multimodal fusion. Previous methods often use techniques such as diffusion models and adversarial learning to reduce the modality gap, but they typically focus on one-to-one…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Sijie Mai , Shiqin Han

Multimodal learning typically relies on the assumption that all modalities are fully available during both the training and inference phases. However, in real-world scenarios, consistently acquiring complete multimodal data presents…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Donggeun Kim , Taesup Kim

Multimodal Federated Learning frequently encounters challenges of client modality heterogeneity, leading to undesired performances for secondary modality in multimodal learning. It is particularly prevalent in audiovisual learning, with…

音频与语音处理 · 电气工程与系统科学 2024-08-29 Tiantian Feng , Tuo Zhang , Salman Avestimehr , Shrikanth S. Narayanan

Vision-language retrieval aims to search for similar instances in one modality based on queries from another modality. The primary objective is to learn cross-modal matching representations in a latent common space. Actually, the assumption…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Yang Yang , Wenjuan Xi , Luping Zhou , Jinhui Tang

Systematics contaminate observables, leading to distribution shifts relative to theoretically simulated signals-posing a major challenge for using pre-trained models to label such observables. Since systematics are often poorly understood…

天体物理仪器与方法 · 物理学 2025-11-18 Sultan Hassan , Sambatra Andrianomena , Benjamin D. Wandelt

Multispectral pedestrian detection has been shown to be effective in improving performance within complex illumination scenarios. However, prevalent double-stream networks in multispectral detection employ two separate feature extraction…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Zizhao Chen , Yeqiang Qian , Xiaoxiao Yang , Chunxiang Wang , Ming Yang

Regularization is used in many different areas of optimization when solutions are sought which not only minimize a given function, but also possess a certain degree of regularity. Popular applications are image denoising, sparse regression…

最优化与控制 · 数学 2021-11-15 Bennet Gebken , Katharina Bieker , Sebastian Peitz

In the surveillance and defense domain, multi-target detection and classification (MTD) is considered essential yet challenging due to heterogeneous inputs from diverse data sources and the computational complexity of algorithms designed…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Ngoc Tuyen Do , Tri Nhu Do

This paper addresses the problem of matching pedestrians across multiple camera views, known as person re-identification. Variations in lighting conditions, environment and pose changes across camera views make re-identification a…

计算机视觉与模式识别 · 计算机科学 2015-12-01 Rahul Rama Varior , Gang Wang

In autonomous driving, environment perception has significantly advanced with the utilization of deep learning techniques for diverse sensors such as cameras, depth sensors, or infrared sensors. The diversity in the sensor stack increases…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Niharika Hegde , Shishir Muralidhara , René Schuster , Didier Stricker

Recent advancements in 3D object detection have benefited from multi-modal information from the multi-view cameras and LiDAR sensors. However, the inherent disparities between the modalities pose substantial challenges. We observe that…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Juhan Cha , Minseok Joo , Jihwan Park , Sanghyeok Lee , Injae Kim , Hyunwoo J. Kim

Detection of surrounding objects and their motion prediction are critical components of a self-driving system. Recently proposed models that jointly address these tasks rely on a number of sensors to achieve state-of-the-art performance.…

机器人学 · 计算机科学 2021-01-12 Abhishek Mohta , Fang-Chieh Chou , Brian C. Becker , Carlos Vallespi-Gonzalez , Nemanja Djuric