English
Related papers

Related papers: A Scalable Framework for Table of Contents Extract…

200 papers

Current abstractive summarization systems outperform their extractive counterparts, but their widespread adoption is inhibited by the inherent lack of interpretability. To achieve the best of both worlds, we propose EASE, an…

Computation and Language · Computer Science 2021-05-17 Haoran Li , Arash Einolghozati , Srinivasan Iyer , Bhargavi Paranjape , Yashar Mehdad , Sonal Gupta , Marjan Ghazvininejad

Electronic theses and dissertations (ETDs) have been proposed, advocated, and generated for more than 25 years. Although ETDs are hosted by commercial or institutional digital library repositories, they are still an understudied type of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Muntabir Hasan Choudhury , Lamia Salsabil , William A. Ingram , Edward A. Fox , Jian Wu

Understanding the semantic meaning of content on the web through the lens of entities and concepts has many practical advantages. However, when building large-scale entity extraction systems, practitioners are facing unique challenges…

Computation and Language · Computer Science 2021-10-04 Xuanting Cai , Quanbin Ma , Pan Li , Jianyu Liu , Qi Zeng , Zhengkan Yang , Pushkar Tripathi

Annually and globally, over three billion radiography examinations and computer tomography scans result in mostly unstructured radiology reports containing free text. Despite the potential benefits of structured reporting, its adoption is…

Computation and Language · Computer Science 2024-06-25 Daniel Reichenpfader , Jonas Knupp , André Sander , Kerstin Denecke

Document generation has gained growing attention in the field of AI-driven content creation. In this work, we push its boundaries by introducing AnyDoc, a framework capable of handling multiple generation tasks across a wide spectrum of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Jiawei Lin , Wanrong Zhu , Vlad I Morariu , Christopher Tensmeyer

Unstructured documents like PDFs contain valuable structured information, but downstream systems require this data in reliable, standardized formats. LLMs are increasingly deployed to automate this extraction, making accuracy and…

Machine Learning · Computer Science 2026-02-17 Nick Ferguson , Josh Pennington , Narek Beghian , Aravind Mohan , Douwe Kiela , Sheshansh Agrawal , Thien Hang Nguyen

Earnings calls are a key source of financial information about public companies. However, extracting information from these calls is difficult. Unlike the templatic filings required by the U.S. Securities and Exchange Commission (SEC) to…

Computation and Language · Computer Science 2026-05-06 Rasmus T. Aavang , Rasmus Tjalk-Bøggild , Alexandre Iolov , Giovanni Rizzi , Mike Zhang , Johannes Bjerva

Understanding documents with rich layouts is an essential step towards information extraction. Business intelligence processes often require the extraction of useful semantic content from documents at a large scale for subsequent…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Sanket Biswas , Ayan Banerjee , Josep Lladós , Umapada Pal

Extractive summarization (ES) aims to generate a concise summary by selecting a subset of sentences from a document while maximizing relevance and minimizing redundancy. Although modern ES systems achieve high accuracy using powerful neural…

Machine Learning · Computer Science 2026-01-26 Ziqing Zeng , Abhimanyu Kumar , Ahmet Efe , Ruihong Yin , Chris H. Kim , Ulya R. Karpuzcu , Sachin S. Sapatnekar

We release EDGAR-CORPUS, a novel corpus comprising annual reports from all the publicly traded companies in the US spanning a period of more than 25 years. To the best of our knowledge, EDGAR-CORPUS is the largest financial NLP corpus…

Computation and Language · Computer Science 2023-05-31 Lefteris Loukas , Manos Fergadiotis , Ion Androutsopoulos , Prodromos Malakasiotis

Chain-of-Thought (CoT) and its variants have markedly advanced the reasoning abilities of Large Language Models (LLMs), yet their monolithic and auto-regressive architecture inherently conflates high-level strategic planning with low-level…

Machine Learning · Computer Science 2025-10-01 Kaisen Yang , Lixuan He , Rushi Shah , Kaicheng Yang , Qinwei Ma , Dianbo Liu , Alex Lamb

Extracting information from academic PDF documents is crucial for numerous indexing, retrieval, and analysis use cases. Choosing the best tool to extract specific content elements is difficult because many, technically diverse tools are…

Information Retrieval · Computer Science 2023-03-20 Norman Meuschke , Apurva Jagdale , Timo Spinde , Jelena Mitrović , Bela Gipp

Questions in open-domain question answering are often ambiguous, allowing multiple interpretations. One approach to handling them is to identify all possible interpretations of the ambiguous question (AQ) and to generate a long-form answer…

Computation and Language · Computer Science 2023-10-24 Gangwoo Kim , Sungdong Kim , Byeongguk Jeon , Joonsuk Park , Jaewoo Kang

Understanding and extracting of information from large documents, such as business opportunities, academic articles, medical documents and technical reports, poses challenges not present in short documents. Such large documents may be…

Computation and Language · Computer Science 2019-10-10 Muhammad Mahbubur Rahman , Tim Finin

Real-world financial filings report critical information about an entity's investment holdings, essential for assessing that entity's risk, profitability, and relationship profile. Yet, these details are often buried in messy, multi-page,…

Artificial Intelligence · Computer Science 2026-02-25 Nicole Cho , Kirsty Fielding , William Watson , Sumitra Ganesh , Manuela Veloso

In this paper, we propose a new system called ASET that allows users to perform structured explorations of text collections in an ad-hoc manner. The main idea of ASET is to use a new two-phase approach that first extracts a superset of…

Computation and Language · Computer Science 2022-03-10 Benjamin Hättasch , Jan-Micha Bodensohn , Carsten Binnig

Maintaining up-to-date, comprehensive documentation for large codebases is a persistent challenge. Recent progress in automated documentation has moved from template-based rules to large language models (LLMs), yet existing tools still…

Software Engineering · Computer Science 2026-04-30 Dong Xu , Mingwei Liu , Xiwen Wang , Jianfeng Zhong , Zibin Zheng

Document-level Event Argument Extraction (EAE) requires the model to extract arguments of multiple events from a single document. Considering the underlying dependencies between these events, recent efforts leverage the idea of "memory",…

Computation and Language · Computer Science 2023-10-26 Quzhe Huang , Yanxi Zhang , Dongyan Zhao

Retrieval-Augmented Generation (RAG) has become the standard approach for grounding large language models in information that was not available during training. While existing datasets and benchmarks focus on web or other public sources,…

Information Retrieval · Computer Science 2026-05-21 Yuhong Sun , Joachim Rahmfeld , Chris Weaver , Weijia Chen , Roshan Desai , Wenxi Huang , Mark H. Butler

In this paper, we study the problem of extracting variable-depth "logical document hierarchy" from long documents, namely organizing the recognized "physical document objects" into hierarchical structures. The discovery of logical document…

Information Retrieval · Computer Science 2021-05-21 Rongyu Cao , Yixuan Cao , Ganbin Zhou , Ping Luo