English
Related papers

Related papers: Fetch-A-Set: A Large-Scale OCR-Free Benchmark for …

200 papers

In AI-facilitated teaching, leveraging various query styles to interpret abstract text descriptions is crucial for ensuring high-quality teaching. However, current retrieval models primarily focus on natural text-image retrieval, making…

Information Retrieval · Computer Science 2025-05-21 Yanhao Jia , Xinyi Wu , Hao Li , Qinglin Zhang , Yuxiao Hu , Shuai Zhao , Wenqi Fan

Aggregation query over free text is a long-standing yet underexplored problem. Unlike ordinary question answering, aggregate queries require exhaustive evidence collection and systems are required to "find all," not merely "find one."…

Artificial Intelligence · Computer Science 2026-02-04 Haojia Zhu , Qinyuan Xu , Haoyu Li , Yuxi Liu , Hanchen Qiu , Jiaoyan Chen , Jiahui Jin

Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for enhancing the capabilities of large language models. However, existing RAG evaluation predominantly focuses on text retrieval and relies on opaque, end-to-end…

Information Retrieval · Computer Science 2025-05-19 Chuan Xu , Qiaosheng Chen , Yutong Feng , Gong Cheng

Legal QA benchmarks have predominantly focused on case law, overlooking the unique challenges of statute-centric regulatory reasoning. In statutory domains, relevant evidence is distributed across hierarchically linked documents, creating a…

Information Retrieval · Computer Science 2026-04-09 Kyubyung Chae , Jewon Yeom , Jeongjae Park , Seunghyun Bae , Ijun Jang , Hyunbin Jin , Jinkwan Jang , Taesup Kim

While semantic segmentation has seen tremendous improvements in the past, there are still significant labeling efforts necessary and the problem of limited generalization to classes that have not been present during training. To address…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Benedikt Blumenstiel , Johannes Jakubik , Hilde Kühne , Michael Vössing

Retrieval Augmented Generation (RAG) has emerged as a standard paradigm for enhancing the factual accuracy and contextual relevance of Large Language Models (LLMs) by integrating retrieval mechanisms. However, existing evaluation frameworks…

Computation and Language · Computer Science 2025-04-11 Mattia Rengo , Senad Beadini , Domenico Alfano , Roberto Abbruzzese

The rapid advancement of multimodal large language models (MLLMs) has profoundly impacted the document domain, creating a wide array of application scenarios. This progress highlights the need for a comprehensive benchmark to evaluate these…

Computation and Language · Computer Science 2025-05-23 Siqi Li , Yufan Shen , Xiangnan Chen , Jiayi Chen , Hengwei Ju , Haodong Duan , Song Mao , Hongbin Zhou , Bo Zhang , Bin Fu , Pinlong Cai , Licheng Wen , Botian Shi , Yong Liu , Xinyu Cai , Yu Qiao

Accurate multi-modal document retrieval is crucial for Retrieval-Augmented Generation (RAG), yet existing benchmarks do not fully capture real-world challenges with their current design. We introduce REAL-MM-RAG, an automatically generated…

Information Retrieval · Computer Science 2025-02-19 Navve Wasserman , Roi Pony , Oshri Naparstek , Adi Raz Goldfarb , Eli Schwartz , Udi Barzelay , Leonid Karlinsky

Extracting hypotheses and their supporting statistical evidence from full-text scientific articles is central to the synthesis of empirical findings, but remains difficult due to document length and the distribution of scientific arguments…

Computation and Language · Computer Science 2026-03-24 Sai Koneru , Jian Wu , Sarah Rajtmajer

Statutory article retrieval (SAR), the task of retrieving statute law articles relevant to a legal question, is a promising application of legal text processing. In particular, high-quality SAR systems can improve the work efficiency of…

Information Retrieval · Computer Science 2023-01-31 Antoine Louis , Gijs van Dijck , Gerasimos Spanakis

Large-scale retrieval systems are often implemented as a cascading sequence of phases -- a first filtering step, in which a large set of candidate documents are extracted using a simple technique such as Boolean matching and/or static…

Information Retrieval · Computer Science 2015-06-03 Charles L. A. Clarke , J. Shane Culpepper , Alistair Moffat

Document retrieval aims at finding the most important documents where a pattern appears in a collection of strings. Traditional pattern-matching techniques yield brute-force document retrieval solutions, which has motivated the research on…

Data Structures and Algorithms · Computer Science 2014-07-02 Gonzalo Navarro , Simon J. Puglisi , Jouni Sirén

This paper addresses the gap between general-purpose text embeddings and the specific demands of item retrieval tasks. We demonstrate the shortcomings of existing models in capturing the nuances necessary for zero-shot performance on item…

Information Retrieval · Computer Science 2024-03-01 Yuxuan Lei , Jianxun Lian , Jing Yao , Mingqi Wu , Defu Lian , Xing Xie

Query-focused summarization (QFS) aims to extract or generate a summary of an input document that directly answers or is relevant to a given query. The lack of large-scale datasets in the form of documents, queries, and summaries has…

Computation and Language · Computer Science 2023-05-23 Ruochen Xu , Song Wang , Yang Liu , Shuohang Wang , Yichong Xu , Dan Iter , Chenguang Zhu , Michael Zeng

Retrieval augmented generation (RAG) has been widely adopted to help Large Language Models (LLMs) to process tasks involving long documents. However, existing retrieval models are not designed for long document retrieval and fail to address…

Information Retrieval · Computer Science 2026-02-13 David Jiahao Fu , Lam Thanh Do , Jiayu Li , Kevin Chen-Chuan Chang

This paper introduces a novel indexing and access method, called Feature- Based Adaptive Tolerance Tree (FATT), using wavelet transform is proposed to organize large image data sets efficiently and to support popular image access mechanisms…

Multimedia · Computer Science 2010-04-09 Dr. P. AnandhaKumar , V. Balamurugan

In this paper, we study the problem of extracting variable-depth "logical document hierarchy" from long documents, namely organizing the recognized "physical document objects" into hierarchical structures. The discovery of logical document…

Information Retrieval · Computer Science 2021-05-21 Rongyu Cao , Yixuan Cao , Ganbin Zhou , Ping Luo

Existing topic modeling and text segmentation methodologies generally require large datasets for training, limiting their capabilities when only small collections of text are available. In this work, we reexamine the inter-related problems…

Information Retrieval · Computer Science 2021-05-26 Qiong Wu , Adam Hare , Sirui Wang , Yuwei Tu , Zhenming Liu , Christopher G. Brinton , Yanhua Li

The Maximum Independent Set problem is fundamental for extracting conflict-free structure from large graphs, with applications in scheduling, recommendation, and network analysis. However, existing heuristics can stagnate when search…

Artificial Intelligence · Computer Science 2025-10-29 Yu Zhang , Witold Pedrycz , Chanjuan Liu , Enqiang Zhu

Efficiently identifying keyphrases that represent a given document is a challenging task. In the last years, plethora of keyword detection approaches were proposed. These approaches can be based on statistical (frequency-based) properties…

Information Retrieval · Computer Science 2023-12-25 Blaž Škrlj , Boshko Koloski , Senja Pollak