English
Related papers

Related papers: Multimodal Information Retrieval for Open World wi…

200 papers

Multimodal IE in social media is difficult because a post may attach multiple images that are weakly related, redundant, or even misleading with respect to the text. In this setting, always-on multimodal fusion wastes computation and can…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Miaobo Hu , Shuhao Hu , Bokun Wang , Rui Chen , Xin Wang , Xiaobo Guo , Daren Zha , Jun Xiao

Multimodal incremental learning needs to digest the information from multiple modalities while concurrently learning new knowledge without forgetting the previously learned information. There are numerous challenges for this task, mainly…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Yi-Lun Lee , Chen-Yu Lee , Wei-Chen Chiu , Yi-Hsuan Tsai

Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories, requiring them to locate relevant evidence across interleaved text and images. Prior…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Aaron Branson Cigres Li , Zhaowei Wang , Yu Zhao , Yiming Du , Haobo Li , Xiyu Ren , Ginny Wong , Simon See , Lishu Luo , Haodong Duan , Pasquale Minervini , Yangqiu Song

Remote sensing cross-modal text-image retrieval (RSCTIR) has gained attention for its utility in information mining. However, challenges remain in effectively integrating global and local information due to variations in remote sensing…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Zengbao Sun , Ming Zhao , Gaorui Liu , André Kaup

Multimodal deep learning systems which employ multiple modalities like text, image, audio, video, etc., are showing better performance in comparison with individual modalities (i.e., unimodal) systems. Multimodal machine learning involves…

Machine Learning · Computer Science 2022-01-19 Anil Rahate , Rahee Walambe , Sheela Ramanna , Ketan Kotecha

In recent years, multimodal multidomain fake news detection has garnered increasing attention. Nevertheless, this direction presents two significant challenges: (1) Failure to Capture Cross-Instance Narrative Consistency: existing models…

Computation and Language · Computer Science 2026-04-30 Yiheng Li , Weihai Lu , Hanyi Yu , Yue Wang

Sparse annotations fundamentally constrain multimodal remote sensing: even recent state-of-the-art supervised methods such as MSFMamba are limited by the availability of labeled data, restricting their practical deployment despite…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Yuzhen Hu , Saurabh Prasad

Multimodal machine learning with missing modalities is an increasingly relevant challenge arising in various applications such as healthcare. This paper extends the current research into missing modalities to the low-data regime, i.e., a…

Machine Learning · Computer Science 2024-03-27 Zhuo Zhi , Ziquan Liu , Moe Elbadawi , Adam Daneshmend , Mine Orlu , Abdul Basit , Andreas Demosthenous , Miguel Rodrigues

Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text-to-image (T2I) tasks, this unified capability allows UMMs to refine outputs after their initial generation, potentially…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Jiayi Guo , Linqing Wang , Jiangshan Wang , Yang Yue , Zeyu Liu , Zhiyuan Zhao , Qinglin Lu , Gao Huang , Chunyu Wang

Multimodal Misinformation Detection (MMD) refers to the task of detecting social media posts involving misinformation, where the post often contains text and image modalities. However, by observing the MMD posts, we hold that the text…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Bing Wang , Ximing Li , Yanjun Wang , Changchun Li , Lin Yuanbo Wu , Buyu Wang , Shengsheng Wang

Textual descriptions for multimodal inputs entail recurrent refinement of queries to produce relevant output images. Despite efforts to address challenges such as scaling model size and data volume, the cost associated with pre-training and…

Machine Learning · Computer Science 2025-08-14 Amit Kumar Jaiswal , Haiming Liu , Ingo Frommholz

Automatic detection of multimodal misinformation has gained a widespread attention recently. However, the potential of powerful Large Language Models (LLMs) for multimodal misinformation detection remains underexplored. Besides, how to…

Computation and Language · Computer Science 2024-04-09 Longzheng Wang , Xiaohan Xu , Lei Zhang , Jiarui Lu , Yongxiu Xu , Hongbo Xu , Minghao Tang , Chuang Zhang

Multimodal knowledge editing represents a critical advancement in enhancing the capabilities of Multimodal Large Language Models (MLLMs). Despite its potential, current benchmarks predominantly focus on coarse-grained knowledge, leaving the…

Computation and Language · Computer Science 2024-02-26 Jiaqi Li , Miaozeng Du , Chuanyi Zhang , Yongrui Chen , Nan Hu , Guilin Qi , Haiyun Jiang , Siyuan Cheng , Bozhong Tian

Joint Multimodal Entity-Relation Extraction (JMERE) is a challenging task that aims to extract entities and their relations from text-image pairs in social media posts. Existing methods for JMERE require large amounts of labeled data.…

Computation and Language · Computer Science 2025-03-25 Li Yuan , Yi Cai , Junsheng Huang

Mutual Information (MI) is often used for feature selection when developing classifier models. Estimating the MI for a subset of features is often intractable. We demonstrate, that under the assumptions of conditional independence, MI…

Machine Learning · Computer Science 2017-06-26 Hemanth Venkateswara , Prasanth Lade , Binbin Lin , Jieping Ye , Sethuraman Panchanathan

Multimodal deep learning has shown strong potential in medical applications by integrating heterogeneous data sources such as medical images and structured clinical variables. However, most existing approaches implicitly assume complete…

Machine Learning · Computer Science 2026-05-13 Camillo Maria Caruso , Valerio Guarrasi , Paolo Soda

The objective of Content-Based Image Retrieval (CBIR) methods is essentially to extract, from large (image) databases, a specified number of images similar in visual and semantic content to a so-called query image. To bridge the semantic…

Information Retrieval · Computer Science 2015-02-12 Smarajit Bose , Amita Pal , Jhimli Mallick , Sunil Kumar , Pratyaydipta Rudra

Neural network representation learning frameworks have recently shown to be highly effective at a wide range of tasks ranging from radiography interpretation via data-driven diagnostics to clinical decision support. This often superior…

Information Retrieval · Computer Science 2018-11-14 Xing Wei , Carsten Eickhoff

Recently, large multimodal models have built a bridge from visual to textual information, but they tend to underperform in remote sensing scenarios. This underperformance is due to the complex distribution of objects and the significant…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Cong Yang , Zuchao Li , Lefei Zhang

Evaluating production-level retrieval systems at scale is a crucial yet challenging task due to the limited availability of a large pool of well-trained human annotators. Large Language Models (LLMs) have the potential to address this…

Information Retrieval · Computer Science 2024-09-19 Kasra Hosseini , Thomas Kober , Josip Krapac , Roland Vollgraf , Weiwei Cheng , Ana Peleteiro Ramallo
‹ Prev 1 3 4 5 6 7 10 Next ›