English
Related papers

Related papers: Modality-Agnostic Attention Fusion for visual sear…

200 papers

Feature selection is essential for high-dimensional biomedical data, enabling stronger predictive performance, reduced computational cost, and improved interpretability in precision medicine applications. Existing approaches face notable…

Machine Learning · Computer Science 2026-01-07 Xiaoyan Sun , Qingyu Meng , Yalu Wen

We propose Adaptive Multi-Style Fusion (AMSF), a reference-based training-free framework that enables controllable fusion of multiple reference styles in diffusion models. Most of the existing reference-based methods are limited by (a)…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Xu Liu , Yibo Lu , Xinxian Wang , Xinyu Wu

The fusion technique is the key to the multimodal emotion recognition task. Recently, cross-modal attention-based fusion methods have demonstrated high performance and strong robustness. However, cross-modal attention suffers from redundant…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Feng Liu , Ziwang Fu , Yunlong Wang , Qijian Zheng

Effective deep feature extraction via feature-level fusion is crucial for multimodal object detection. However, previous studies often involve complex training processes that integrate modality-specific features by stacking multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Lei Hao , Lina Xu , Chang Liu , Yanni Dong

We study few-shot learning in natural language domains. Compared to many existing works that apply either metric-based or optimization-based meta-learning to image domain with low inter-task variance, we consider a more realistic setting,…

Computation and Language · Computer Science 2018-05-22 Mo Yu , Xiaoxiao Guo , Jinfeng Yi , Shiyu Chang , Saloni Potdar , Yu Cheng , Gerald Tesauro , Haoyu Wang , Bowen Zhou

Image modality is not perfect as it often fails in certain conditions, e.g., night and fast motion. This significantly limits the robustness and versatility of existing multi-modal (i.e., Image+X) semantic segmentation methods when…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Xu Zheng , Yuanhuiyi Lyu , Lin Wang

Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Xunpeng Yi , Han Xu , Hao Zhang , Linfeng Tang , Jiayi Ma

A critical challenge to image-text retrieval is how to learn accurate correspondences between images and texts. Most existing methods mainly focus on coarse-grained correspondences based on co-occurrences of semantic objects, while failing…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Guoliang Wang , Yanlei Shang , Yong Chen

Multimodal recommendation systems are increasingly becoming foundational technologies for e-commerce and content platforms, enabling personalized services by jointly modeling users' historical behaviors and the multimodal features of items…

Information Retrieval · Computer Science 2025-09-12 Kelin Ren , Chan-Yang Ju , Dong-Ho Lee

Multi-modal learning has shown exceptional performance in various tasks, especially in medical applications, where it integrates diverse medical information for comprehensive diagnostic evidence. However, there still are several challenges…

Machine Learning · Computer Science 2024-11-19 Lin Fan , Yafei Ou , Cenyang Zheng , Pengyu Dai , Tamotsu Kamishima , Masayuki Ikebe , Kenji Suzuki , Xun Gong

Heterogeneous face recognition is a challenging task due to the large modality discrepancy and insufficient cross-modal samples. Most existing works focus on discriminative feature transformation, metric learning and cross-modal face…

Computer Vision and Pattern Recognition · Computer Science 2020-08-11 Yingguo Xu , Lei Zhang , Qingyan Duan

Multimodal large language models (MLLMs) have shown impressive capabilities, yet they often struggle to effectively capture the fine-grained textual information within images crucial for accurate image translation. This often leads to a…

Computation and Language · Computer Science 2026-04-21 Bo Li , Ningyuan Deng , Tianyu Dong , Shaobo Wang , Shaolin Zhu , Lijie Wen

Nowadays, cross-modal retrieval plays an indispensable role to flexibly find information across different modalities of data. Effectively measuring the similarity between different modalities of data is the key of cross-modal retrieval.…

Computer Vision and Pattern Recognition · Computer Science 2017-08-17 Yuxin Peng , Jinwei Qi , Yuxin Yuan

Multi-modal learning has emerged as a crucial research direction, as integrating textual and visual information can substantially enhance performance in tasks such as classification, retrieval, and scene understanding. Despite advances with…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Md. Mithun Hossain , Md. Shakil Hossain , Sudipto Chaki , M. F. Mridha

Visual question answering (VQA) is challenging because it requires a simultaneous understanding of both the visual content of images and the textual content of questions. The approaches used to represent the images and questions in a…

Computer Vision and Pattern Recognition · Computer Science 2017-08-07 Zhou Yu , Jun Yu , Jianping Fan , Dacheng Tao

Multi-modality image fusion (MMIF) aims to integrate complementary information from different modalities into a single fused image to represent the imaging scene and facilitate downstream visual tasks comprehensively. In recent years,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Zhe Li , Haiwei Pan , Kejia Zhang , Yuhua Wang , Fengming Yu

Image-text matching aims to find matched cross-modal pairs accurately. While current methods often rely on projecting cross-modal features into a common embedding space, they frequently suffer from imbalanced feature representations across…

Information Retrieval · Computer Science 2024-01-19 Zuhui Wang , Yunting Yin , I. V. Ramakrishnan

Most few-shot learning models utilize only one modality of data. We would like to investigate qualitatively and quantitatively how much will the model improve if we add an extra modality (i.e. text description of the image), and how it…

Computer Vision and Pattern Recognition · Computer Science 2021-07-27 Zilun Zhang , Shihao Ma , Yichun Zhang

Action quality assessment (AQA) is to assess how well an action is performed. Previous works perform modelling by only the use of visual information, ignoring audio information. We argue that although AQA is highly dependent on visual…

Signal Processing · Electrical Eng. & Systems 2025-03-06 Ling-An Zeng , Wei-Shi Zheng

In many real-world scenarios, acquiring all features of a data instance can be expensive or impractical due to monetary cost, latency, or privacy concerns. Active Feature Acquisition (AFA) addresses this challenge by dynamically selecting a…

Machine Learning · Computer Science 2026-02-24 Valter Schütz , Han Wu , Reza Rezvan , Linus Aronsson , Morteza Haghir Chehreghani
‹ Prev 1 4 5 6 7 8 10 Next ›