English
Related papers

Related papers: MMDocIR: Benchmarking Multimodal Retrieval for Lon…

200 papers

Multimodal information extraction (MIE) is crucial for scientific literature, where valuable data is often spread across text, figures, and tables. In materials science, extracting structured information from research articles can…

Computation and Language · Computer Science 2024-10-29 Ghazal Khalighinejad , Sharon Scott , Ollie Liu , Kelly L. Anderson , Rickard Stureborg , Aman Tyagi , Bhuwan Dhingra

The rapid advancement of Multi-modal Large Language Models (MLLMs) has expanded their capabilities beyond high-level vision tasks. Nevertheless, their potential for Document Image Quality Assessment (DIQA) remains underexplored. To bridge…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Jiaxi Huang , Dongxu Wu , Hanwei Zhu , Lingyu Zhu , Jun Xing , Xu Wang , Baoliang Chen

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their…

Source attribution aims to enhance the reliability of AI-generated answers by including references for each statement, helping users validate the provided answers. However, existing work has primarily focused on text-only scenario and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Seokwon Song , Minsu Park , Gunhee Kim

Despite significant progress in multimodal large language models (MLLMs), their performance on complex, multi-page document comprehension remains inadequate, largely due to the lack of high-quality, document-level datasets. While current…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yuchen Duan , Zhe Chen , Yusong Hu , Weiyun Wang , Shenglong Ye , Botian Shi , Lewei Lu , Qibin Hou , Tong Lu , Hongsheng Li , Jifeng Dai , Wenhai Wang

Image-text retrieval, as a fundamental and important branch of information retrieval, has attracted extensive research attentions. The main challenge of this task is cross-modal semantic understanding and matching. Some recent works focus…

Computer Vision and Pattern Recognition · Computer Science 2023-04-24 Weijing Chen , Linli Yao , Qin Jin

Documents are visually rich structures that convey information through text, but also figures, page layouts, tables, or even fonts. Since modern retrieval systems mainly rely on the textual information they extract from document pages to…

Information Retrieval · Computer Science 2025-03-03 Manuel Faysse , Hugues Sibille , Tony Wu , Bilel Omrani , Gautier Viaud , Céline Hudelot , Pierre Colombo

Accurate multi-modal document retrieval is crucial for Retrieval-Augmented Generation (RAG), yet existing benchmarks do not fully capture real-world challenges with their current design. We introduce REAL-MM-RAG, an automatically generated…

Information Retrieval · Computer Science 2025-02-19 Navve Wasserman , Roi Pony , Oshri Naparstek , Adi Raz Goldfarb , Eli Schwartz , Udi Barzelay , Leonid Karlinsky

Existing multimodal browsing benchmarks often fail to require genuine multimodal reasoning, as many tasks can be solved with text-only heuristics without vision-in-the-loop verification. We introduce MMSearch-Plus, a 311-task benchmark that…

Artificial Intelligence · Computer Science 2026-03-20 Xijia Tao , Yihua Teng , Xinxing Su , Xinyu Fu , Jihao Wu , Chaofan Tao , Ziru Liu , Haoli Bai , Rui Liu , Lingpeng Kong

Large Vision-Language Models (LVLMs) have demonstrated strong multimodal reasoning capabilities on long and complex documents. However, their high memory footprint makes them impractical for deployment on resource-constrained edge devices.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Tanveer Hannan , Dimitrios Mallios , Parth Pathak , Faegheh Sardari , Thomas Seidl , Gedas Bertasius , Mohsen Fayyaz , Sunando Sengupta

Current multimodal information retrieval studies mainly focus on single-image inputs, which limits real-world applications involving multiple images and text-image interleaved content. In this work, we introduce the text-image interleaved…

Computation and Language · Computer Science 2025-02-19 Xin Zhang , Ziqi Dai , Yongqi Li , Yanzhao Zhang , Dingkun Long , Pengjun Xie , Meishan Zhang , Jun Yu , Wenjie Li , Min Zhang

We present a new state-of-the-art on the text to video retrieval task on MSRVTT and LSMDC benchmarks where our model outperforms all previous solutions by a large margin. Moreover, state-of-the-art results are achieved with a single model…

Computer Vision and Pattern Recognition · Computer Science 2021-11-09 Maksim Dzabraev , Maksim Kalashnikov , Stepan Komkov , Aleksandr Petiushko

Existing Multimodal Large Language Models (MLLMs) are predominantly trained and tested on consistent visual-textual inputs, leaving open the question of whether they can handle inconsistencies in real-world, layout-rich content. To bridge…

Computation and Language · Computer Science 2025-06-12 Qianqi Yan , Yue Fan , Hongquan Li , Shan Jiang , Yang Zhao , Xinze Guan , Ching-Chen Kuo , Xin Eric Wang

Information Retrieval (IR) methods aim to identify documents relevant to a query, which have been widely applied in various natural language tasks. However, existing approaches typically consider only the textual content within documents,…

Computation and Language · Computer Science 2026-01-26 Jaewoo Lee , Joonho Ko , Jinheon Baek , Soyeong Jeong , Sung Ju Hwang

Recent advances in Retrieval-Augmented Generation (RAG) have significantly improved response accuracy and relevance by incorporating external knowledge into Large Language Models (LLMs). However, existing RAG methods primarily focus on…

Machine Learning · Computer Science 2025-04-22 Qinhan Yu , Zhiyou Xiao , Binghui Li , Zhengren Wang , Chong Chen , Wentao Zhang

In real-world documents, the information relevant to a user query may reside anywhere from the beginning to the end. This makes position bias -- a systematic tendency of retrieval models to favor or neglect content based on its location --…

Information Retrieval · Computer Science 2026-03-13 Ziyang Zeng , Dun Zhang , Yu Yan , Xu Sun , Cuiqiaoshu Pan , Yudong Zhou , Yuqing Yang

Document translation poses a challenge for Neural Machine Translation (NMT) systems. Most document-level NMT systems rely on meticulously curated sentence-level parallel data, assuming flawless extraction of text from documents along with…

Computation and Language · Computer Science 2024-06-13 Benjamin Hsu , Xiaoyu Liu , Huayang Li , Yoshinari Fujinuma , Maria Nadejde , Xing Niu , Yair Kittenplon , Ron Litman , Raghavendra Pappagari

Multimodal learning is a recent challenge that extends unimodal learning by generalizing its domain to diverse modalities, such as texts, images, or speech. This extension requires models to process and relate information from multiple…

Information Retrieval · Computer Science 2022-09-29 Cheng-An Hsieh , Cheng-Ping Hsieh , Pu-Jen Cheng

We introduce MMTR-Bench, a benchmark designed to evaluate the intrinsic ability of Multimodal Large Language Models (MLLMs) to reconstruct masked text directly from visual context. Unlike conventional question-answering tasks, MMTR-Bench…

Artificial Intelligence · Computer Science 2026-04-28 Jindi Guo , Chaozheng Huang , Xi Fang

We consider and propose a new problem of retrieving audio files relevant to multimodal design document inputs comprising both textual elements and visual imagery, e.g., birthday/greeting cards. In addition to enhancing user experience,…

Multimedia · Computer Science 2023-03-01 Prachi Singh , Srikrishna Karanam , Sumit Shekhar
‹ Prev 1 3 4 5 6 7 10 Next ›