中文
相关论文

相关论文: Modality Dominance-Aware Optimization for Embodied…

200 篇论文

In multimodal learning, dominant modalities often overshadow others, limiting generalization. We propose Modality-Aware Sharpness-Aware Minimization (M-SAM), a model-agnostic framework that applies to many modalities and supports early and…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Hossein R. Nowdeh , Jie Ji , Xiaolong Ma , Fatemeh Afghah

Most existing multimodal trackers adopt uniform fusion strategies, overlooking the inherent differences between modalities. Moreover, they propagate temporal information through mixed tokens, leading to entangled and less discriminative…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Shilei Wang , Pujian Lai , Dong Gao , Jifeng Ning , Gong Cheng

Multimodal deep learning (MDL) has emerged as a transformative approach in computational pathology. By integrating complementary information from multiple data sources, MDL models have demonstrated superior predictive performance across…

定量方法 · 定量生物学 2025-11-17 Seth Alain Chang , Muhammad Mueez Amjad , Noorul Wahab , Ethar Alzaid , Nasir Rajpoot , Adam Shephard

Visual place classification from a first-person-view monocular RGB image is a fundamental problem in long-term robot navigation. A difficulty arises from the fact that RGB image classifiers are often vulnerable to spatial and appearance…

计算机视觉与模式识别 · 计算机科学 2023-05-12 Tomoya Iwasaki , Kanji Tanaka , Kenta Tsukahara

Multispectral object detection, utilizing both visible (RGB) and thermal infrared (T) modals, has garnered significant attention for its robust performance across diverse weather and lighting conditions. However, effectively exploiting the…

计算机视觉与模式识别 · 计算机科学 2024-05-30 Jinzhong Wang , Xuetao Tian , Shun Dai , Tao Zhuo , Haorui Zeng , Hongjuan Liu , Jiaqi Liu , Xiuwei Zhang , Yanning Zhang

Image outpainting technology generates visually plausible content regardless of authenticity, making it unreliable to be applied in practice. Thus, we propose a reliable image outpainting task, introducing the sparse depth from LiDARs to…

计算机视觉与模式识别 · 计算机科学 2023-02-17 Lei Zhang , Kang Liao , Chunyu Lin , Yao Zhao

Human action recognition (HAR) with multi-modal inputs (RGB-D, skeleton, point cloud) can achieve high accuracy but typically relies on large labeled datasets and degrades sharply when sensors fail or are noisy. We present Robust…

信号处理 · 电气工程与系统科学 2025-11-18 Hasan Akgul , Mari Eplik , Javier Rojas , Akira Yamamoto , Rajesh Kumar , Maya Singh

Multimodal industrial surface defect detection (MISDD) aims to identify and locate defect in industrial products by fusing RGB and 3D modalities. This article focuses on modality-missing problems caused by uncertain sensors availability in…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Shuai Jiang , Yunfeng Ma , Jingyu Zhou , Yuan Bian , Yaonan Wang , Min Liu

Multi-modality imaging is widely used in clinical practice and biomedical research to gain a comprehensive understanding of an imaging subject. Currently, multi-modality imaging is accomplished by post hoc fusion of independently…

图像与视频处理 · 电气工程与系统科学 2024-10-01 Lingting Zhu , Yizheng Chen , Lianli Liu , Lei Xing , Lequan Yu

Multi-modal ophthalmic image classification plays a key role in diagnosing eye diseases, as it integrates information from different sources to complement their respective performances. However, recent improvements have mainly focused on…

图像与视频处理 · 电气工程与系统科学 2024-05-29 Ke Zou , Tian Lin , Zongbo Han , Meng Wang , Xuedong Yuan , Haoyu Chen , Changqing Zhang , Xiaojing Shen , Huazhu Fu

Robust perception at night remains challenging for thermal-infrared detection: low contrast and weak high-frequency cues lead to duplicate, overlapping boxes, missed small objects, and class confusion. Prior remedies either translate TIR to…

计算机视觉与模式识别 · 计算机科学 2025-11-04 SiWoo Kim , JhongHyun An

Understanding dark scenes based on multi-modal image data is challenging, as both the visible and auxiliary modalities provide limited semantic information for the task. Previous methods focus on fusing the two modalities but neglect the…

计算机视觉与模式识别 · 计算机科学 2023-11-22 Xiaoyu Dong , Naoto Yokoya

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a diverse range of multimodal tasks. However, these models suffer from a core problem known as text dominance: they depend heavily on text for their…

计算与语言 · 计算机科学 2025-08-15 Huyu Wu , Meng Tang , Xinhan Zheng , Haiyun Jiang

In this work, we propose a new generic multi-modality domain adaptation framework called Progressive Modality Cooperation (PMC) to transfer the knowledge learned from the source domain to the target domain by exploiting multiple modality…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Weichen Zhang , Dong Xu , Jing Zhang , Wanli Ouyang

Infrared-visible object detection improves detection performance by combining complementary features from multispectral images. Existing backbone-specific and backbone-shared approaches still suffer from the problems of severe bias of…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Chunjin Yang , Xiwei Zhang , Yiming Xiao , Fanman Meng

Multimodal deep learning has shown strong potential in medical applications by integrating heterogeneous data sources such as medical images and structured clinical variables. However, most existing approaches implicitly assume complete…

机器学习 · 计算机科学 2026-05-13 Camillo Maria Caruso , Valerio Guarrasi , Paolo Soda

Multi-modal fusion methods often suffer from two types of representation collapse: feature collapse where individual dimensions lose their discriminative power (as measured by eigenspectra), and modality collapse where one dominant modality…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Seulgi Kim , Kiran Kokilepersaud , Mohit Prabhushankar , Ghassan AlRegib

Visual tracking often faces challenges such as invalid targets and decreased performance in low-light conditions when relying solely on RGB image sequences. While incorporating additional modalities like depth and infrared data has proven…

计算机视觉与模式识别 · 计算机科学 2023-12-25 Lei Liu , Mengya Zhang , Cheng Li , Chenglong Li , Jin Tang

Industrial anomaly detection based on RGB-3D multimodal data has emerged as a mainstream paradigm for intelligent quality inspection. However, existing unsupervised methods suffer from two critical limitations: ambiguous cross-modal…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Zewen Li , Shuo Ye , Zitong Yu , Weicheng Xie , Linlin Shen

Multimodal sensing has proven valuable for visual tracking, as different sensor types offer unique strengths in handling one specific challenging scene where object appearance varies. While a generalist model capable of leveraging all…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Yuedong Tan , Zongwei Wu , Yuqian Fu , Zhuyun Zhou , Guolei Sun , Eduard Zamfi , Chao Ma , Danda Pani Paudel , Luc Van Gool , Radu Timofte