English
Related papers

Related papers: SIMARA: a database for key-value information extra…

200 papers

Despite increasing interest in Syriac studies and growing digital availability of Syriac texts, there is currently no up-to-date infrastructure for discovering, identifying, classifying, and referencing works of Syriac literature. The…

Digital Libraries · Computer Science 2021-07-01 Nathan P. Gibson , David A. Michelson , Daniel L. Schwartz

Indexing learning documents using the Learning Object Metadata (LOM) is often carried out manually by archivists. Filling out the LOM fields is a long and difficult task, requiring a complete reading and a full knowledge on the topic dealt…

Information Retrieval · Computer Science 2016-11-27 Carlo Abi Chahine , Jean-Philippe Kotowicz , Nathalie Chaignaud , Jean-Pierre Pécuchet

Generative retrieval (Wang et al., 2022; Tay et al., 2022) is a popular approach for end-to-end document retrieval that directly generates document identifiers given an input query. We introduce summarization-based document IDs, in which…

Computation and Language · Computer Science 2024-10-31 Haoxin Li , Daniel Cheng , Phillip Keung , Jungo Kasai , Noah A. Smith

Existing scholarly information extraction (SIE) datasets focus on scientific papers and overlook implementation-level details in code repositories. README files describe datasets, source code, and other implementation-level artifacts,…

Computation and Language · Computer Science 2026-03-09 Genet Asefa Gesese , Zongxiong Chen , Shufan Jiang , Mary Ann Tan , Zhaotai Liu , Sonja Schimmler , Harald Sack

Maps are an important source of information in archaeology and other sciences. Users want to search for historical maps to determine recorded history of the political geography of regions at different eras, to find out where exactly…

Digital Libraries · Computer Science 2009-01-27 Qingzhao Tan , Prasenjit Mitra , C. Lee Giles

Automatic terminology processing appeared 10 years ago when electronic corpora became widely available. Such processing may be statistically or linguistically based and produces terminology resources that can be used in a number of…

Computers and Society · Computer Science 2014-12-16 C. Enguehard , B. Daille , E. Morin

Extracting key information from documents represents a large portion of business workloads and therefore offers a high potential for efficiency improvements and process automation. With recent advances in Deep Learning, a plethora of Deep…

Information Retrieval · Computer Science 2025-07-21 Alexander Michael Rombach , Peter Fettke

Generative retrieval, which is a new advanced paradigm for document retrieval, has recently attracted research interests, since it encodes all documents into the model and directly generates the retrieved documents. However, its power is…

Information Retrieval · Computer Science 2023-10-31 Tianchi Yang , Minghui Song , Zihan Zhang , Haizhen Huang , Weiwei Deng , Feng Sun , Qi Zhang

The scientific literature is growing exponentially, and professionals are no more able to cope with the current amount of publications. Text mining provided in the past methods to retrieve and extract information from text; however, most of…

Computation and Language · Computer Science 2019-02-27 Nikola Milosevic , Cassie Gregson , Robert Hernandez , Goran Nenadic

Methods for linking individuals across historical data sets, typically in combination with AI based transcription models, are developing rapidly. Probably the single most important identifier for linking is personal names. However, personal…

Computer Vision and Pattern Recognition · Computer Science 2022-03-11 Christian M. Dahl , Torben Johansen , Emil N. Sørensen , Simon Wittrock

Many documents, that we call templatized documents, are programmatically generated by populating fields in a visual template. Effective data extraction from these documents is crucial to supporting downstream analytical tasks. Current data…

Databases · Computer Science 2025-01-14 Yiming Lin , Mawil Hasan , Rohan Kosalge , Alvin Cheung , Aditya G. Parameswaran

We propose Referral-Augmented Retrieval (RAR), a simple technique that concatenates document indices with referrals, i.e. text from other documents that cite or link to the given document, to provide significant performance gains for…

Computation and Language · Computer Science 2023-05-25 Michael Tang , Shunyu Yao , John Yang , Karthik Narasimhan

Cross-lingual information retrieval (CLIR) helps users find documents in languages different from their queries. This is especially important in academic search, where key research is often published in non-English languages. We present…

Information Retrieval · Computer Science 2025-11-20 Francisco Valentini , Diego Kozlowski , Vincent Larivière

The CENDARI infrastructure is a research supporting platform designed to provide tools for transnational historical research, focusing on two topics: Medieval culture and World War I. It exposes to the end users modern web-based tools…

The extraction of key information from receipts is a complex task that involves the recognition and extraction of text from scanned receipts. This process is crucial as it enables the retrieval of essential content and organizing it into…

Computation and Language · Computer Science 2024-03-27 Abdelrahman Abdallah , Mahmoud Abdalla , Mohamed Elkasaby , Yasser Elbendary , Adam Jatowt

When extracting information from handwritten documents, text transcription and named entity recognition are usually faced as separate subsequent tasks. This has the disadvantage that errors in the first module affect heavily the performance…

Computer Vision and Pattern Recognition · Computer Science 2018-03-23 Manuel Carbonell , Mauricio Villegas , Alicia Fornés , Josep Lladós

A large amount of data is present on the web. It contains huge number of web pages and to find suitable information from them is very cumbersome task. There is need to organize data in formal manner so that user can easily access and use…

Information Retrieval · Computer Science 2014-03-28 Gagandeep Singh , Vishal Jain

Humanitarian action is accompanied by a mass of reports, summaries, news, and other documents. To guide its activities, important information must be quickly extracted from such free-text resources. Quantities, such as the number of people…

Computation and Language · Computer Science 2024-08-12 Daniele Liberatore , Kyriaki Kalimeri , Derya Sever , Yelena Mejova

The plethora of digitalised historical document datasets released in recent years has rekindled interest in advancing the field of handwriting pattern recognition. In the same vein, a recently published data set, known as ARDIS, presents…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Mengqiao Zhao , Andre G. Hochuli , Abbas Cheddad

In this paper we report on the initial activities carried out within a collaboration between Consob and Sapienza University. We focus on Information Extraction from documents describing financial instruments. We discuss how we automate this…

Computation and Language · Computer Science 2022-02-03 Domenico Lembo , Alessandra Limosani , Francesca Medda , Alessandra Monaco , Federico Maria Scafoglieri