English
Related papers

Related papers: Learning Emergent Modular Representations in Multi…

200 papers

Multimodal human action understanding is a significant problem in computer vision, with the central challenge being the effective utilization of the complementarity among diverse modalities while maintaining model efficiency. However, most…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Hongsong Wang , Heng Fei , Bingxuan Dai , Jie Gui

Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Zhu Liu , Jinyuan Liu , Guanyao Wu , Long Ma , Xin Fan , Risheng Liu

Healthcare systems generate diverse multimodal data, including Electronic Health Records (EHR), clinical notes, and medical images. Effectively leveraging this data for clinical prediction is challenging, particularly as real-world samples…

Machine Learning · Computer Science 2025-09-01 Xiaoyang Wang , Christopher C. Yang

Bird's Eye View (BEV) perception systems based on multi-sensor feature fusion have become a fundamental cornerstone for end-to-end autonomous driving. However, existing multi-modal BEV methods commonly suffer from limited input…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Qi Xiang , Kunsong Shi , Zhigui Lin , Lei He

Anomaly detection in medical imaging is essential for identifying rare pathological conditions, particularly when annotated abnormal samples are limited. We propose a hybrid anomaly detection framework that integrates self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Pritam Kar , Gouri Lakshmi S , Saptarshi Bej

Unified Multimodal Models (UMMs) built on shared autoregressive (AR) transformers are attractive for their architectural simplicity. However, we identify a critical limitation: when trained on multimodal inputs, modality-shared transformers…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Jitai Hao , Hao Liu , Xinyan Xiao , Qiang Huang , Jun Yu

The rapid development of diagnostic technologies in healthcare is leading to higher requirements for physicians to handle and integrate the heterogeneous, yet complementary data that are produced during routine practice. For instance, the…

Machine Learning · Computer Science 2023-01-30 Can Cui , Haichun Yang , Yaohong Wang , Shilin Zhao , Zuhayr Asad , Lori A. Coburn , Keith T. Wilson , Bennett A. Landman , Yuankai Huo

Ophthalmic images may contain identical-looking pathologies that can cause failure in automated techniques to distinguish different retinal degenerative diseases. Additionally, reliance on large annotated datasets and lack of knowledge…

Image and Video Processing · Electrical Eng. & Systems 2022-08-02 Sharif Amit Kamran , Khondker Fariha Hossain , Alireza Tavakkoli , Stewart Lee Zuckerbrod , Salah A. Baker

For the visible-infrared person re-identification (VIReID) task, one of the major challenges is the modality gaps between visible (VIS) and infrared (IR) images. However, the training samples are usually limited, while the modality gaps are…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Yukang Zhang , Hanzi Wang

Real-time semantic segmentation plays an important role in practical applications such as self-driving and robots. Most semantic segmentation research focuses on improving estimation accuracy with little consideration on efficiency. Several…

Computer Vision and Pattern Recognition · Computer Science 2020-01-01 Shao-Yuan Lo , Hsueh-Ming Hang , Sheng-Wei Chan , Jing-Jhih Lin

One of the early weaknesses identified in deep neural networks trained for image classification tasks was their inability to provide low confidence predictions on out-of-distribution (OOD) data that was significantly different from the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Evelyn Mannix , Howard Bondell

Recently, deep-learning-based approaches have been widely studied for deformable image registration task. However, most efforts directly map the composite image representation to spatial transformation through the convolutional neural…

Image and Video Processing · Electrical Eng. & Systems 2022-07-08 Jiashun Chen , Donghuan Lu , Yu Zhang , Dong Wei , Munan Ning , Xinyu Shi , Zhe Xu , Yefeng Zheng

Modular Neural Networks (MNNs) demonstrate various advantages over monolithic models. Existing MNNs are generally $\textit{explicit}$: their modular architectures are pre-defined, with individual modules expected to implement distinct…

Machine Learning · Computer Science 2024-04-02 Zihan Qiu , Zeyu Huang , Jie Fu

Osteoporosis is a common condition that increases fracture risk, especially in older adults. Early diagnosis is vital for preventing fractures, reducing treatment costs, and preserving mobility. However, healthcare providers face challenges…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Mehdi Hosseini Chagahi , Saeed Mohammadi Dashtaki , Niloufar Delfan , Nadia Mohammadi , Farshid Rostami Pouria , Behzad Moshiri , Md. Jalil Piran , Oliver Faust

We present a model that can perform multiple vision tasks and can be adapted to other downstream tasks efficiently. Despite considerable progress in multi-task learning, most efforts focus on learning from multi-label data: a single image…

Computer Vision and Pattern Recognition · Computer Science 2023-06-30 Zitian Chen , Mingyu Ding , Yikang Shen , Wei Zhan , Masayoshi Tomizuka , Erik Learned-Miller , Chuang Gan

Medical generative models, acknowledged for their high-quality sample generation ability, have accelerated the fast growth of medical applications. However, recent works concentrate on separate medical generation models for distinct medical…

Image and Video Processing · Electrical Eng. & Systems 2024-03-08 Chenlu Zhan , Yu Lin , Gaoang Wang , Hongwei Wang , Jian Wu

Depression is a severe mental disorder, and reliable identification plays a critical role in early intervention and treatment. Multimodal depression detection aims to improve diagnostic performance by jointly modeling complementary…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Chongxiao Wang , Junjie Liang , Peng Cao , Jinzhu Yang , Osmar R. Zaiane

Medical diagnostic applications require models that can process multimodal medical inputs (images, patient histories, lab results) and generate diverse outputs including both textual reports and visual content (annotations, segmentation…

Deep learning models for medical data are typically trained using task specific objectives that encourage representations to collapse onto a small number of discriminative directions. While effective for individual prediction problems, this…

Machine Learning · Computer Science 2026-02-10 Yuanyun Zhang , Mingxuan Zhang , Siyuan Li , Zihan Wang , Haoran Chen , Wenbo Zhou , Shi Li

Emergency department triage relies heavily on both quantitative vital signs and qualitative clinical notes, yet multimodal machine learning models predicting triage acuity often suffer from modality collapse by over-relying on structured…

Machine Learning · Computer Science 2026-04-14 Tyler Yang , Romal Mitr