English
Related papers

Related papers: SIMARA: a database for key-value information extra…

200 papers

A hidden database refers to a dataset that an organization makes accessible on the web by allowing users to issue queries through a search interface. In other words, data acquisition from such a source is not by following static…

Databases · Computer Science 2012-08-02 Cheng Sheng , Nan Zhang , Yufei Tao , Xin Jin

In this paper, we present a pipeline for image extraction from historical documents using foundation models, and evaluate text-image prompts and their effectiveness on humanities datasets of varying levels of complexity. The motivation for…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Hassan El-Hajj , Matteo Valleriani

Archaeological sites are the physical remains of past human activity and one of the main sources of information about past societies and cultures. However, they are also the target of malevolent human actions, especially in countries having…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Elliot Vincent , Mehraïl Saroufim , Jonathan Chemla , Yves Ubelmann , Philippe Marquis , Jean Ponce , Mathieu Aubry

Information extraction from handwritten documents involves traditionally three distinct steps: Document Layout Analysis, Handwritten Text Recognition, and Named Entity Recognition. Recent approaches have attempted to integrate these steps…

Artificial Intelligence · Computer Science 2026-02-03 Thomas Constum , Pierrick Tranouez , Thierry Paquet

This paper is a survey discussing Information Retrieval concepts, methods, and applications. It goes deep into the document and query modelling involved in IR systems, in addition to pre-processing operations such as removing stop words and…

Information Retrieval · Computer Science 2012-12-11 Youssef Bassil

Offline Handwritten Text Recognition (HTR) systems play a crucial role in applications such as historical document digitization, automatic form processing, and biometric authentication. However, their performance is often hindered by the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Yassin Hussein Rassul , Aram M. Ahmed , Polla Fattah , Bryar A. Hassan , Arwaa W. Abdulkareem , Tarik A. Rashid , Joan Lu

Automatically summarizing large text collections is a valuable tool for document research, with applications in journalism, academic research, legal work, and many other fields. In this work, we contrast two classes of systems for…

Computation and Language · Computer Science 2025-02-11 Adithya Pratapa , Teruko Mitamura

Current approaches to automatic summarization of scientific papers generate informative summaries in the form of abstracts. However, abstracts are not intended to show the relationship between a paper and the references cited in it. We…

Computation and Language · Computer Science 2023-11-14 Shahbaz Syed , Ahmad Dawar Hakimi , Khalid Al-Khatib , Martin Potthast

Multimodal approaches have shown great promise for searching and navigating digital collections held by libraries, archives, and museums. In this paper, we introduce map-RAS: a retrieval-augmented search system for historic maps. In…

Information Retrieval · Computer Science 2025-10-30 Jamie Mahowald , Benjamin Charles Germain Lee

Despite biographies are widely spread within the Semantic Web, resources and approaches to automatically extract biographical events are limited. Such limitation reduces the amount of structured, machine-readable biographical information,…

Computation and Language · Computer Science 2022-06-09 Marco Antonio Stranisci , Enrico Mensa , Ousmane Diakite , Daniele Radicioni , Rossana Damiano

We study the problem of deep recall model in industrial web search, which is, given a user query, retrieve hundreds of most relevance documents from billions of candidates. The common framework is to train two encoding models based on…

Information Retrieval · Computer Science 2020-07-06 Yusi Zhang , Chuanjie Liu , Angen Luo , Hui Xue , Xuan Shan , Yuxiang Luo , Yiqian Xia , Yuanchi Yan , Haidong Wang

The automation of document processing is gaining recent attention due to the great potential to reduce manual work through improved methods and hardware. Neural networks have been successfully applied before - even though they have been…

Computation and Language · Computer Science 2021-06-15 Martin Holeček

The preservation of cultural heritage is increasingly transitioning towards data-driven predictive maintenance and "Digital Twin" construction. However, the mechanical constitutive models required for high-fidelity simulations remain…

Databases · Computer Science 2026-02-19 Rui Hu , Yue Wu , Tianhao Su , Yin Wang , Shunbo Hu , Jizhong Huang

We present a corpus of 5,000 richly annotated abstracts of medical articles describing clinical randomized controlled trials. Annotations include demarcations of text spans that describe the Patient population enrolled, the Interventions…

Computation and Language · Computer Science 2018-06-13 Benjamin Nye , Junyi Jessy Li , Roma Patel , Yinfei Yang , Iain J. Marshall , Ani Nenkova , Byron C. Wallace

SYNTAGMA is a rule-based parsing system, structured on two levels: a general parsing engine and a language specific grammar. The parsing engine is a language independent program, while grammar and language specific rules and resources are…

Computation and Language · Computer Science 2016-01-22 Daniel Christen

The introduction of computerized medical records in hospitals has reduced burdensome activities like manual writing and information fetching. However, the data contained in medical records are still far underutilized, primarily because…

A critical aspect of power systems research is the availability of suitable data, access to which is limited by privacy concerns and the sensitive nature of energy infrastructure. This lack of data, in turn, hinders the development of…

Machine Learning · Computer Science 2021-10-27 Minas Chatzos , Mathieu Tanneau , Pascal Van Hentenryck

Researchers continually perform corroborative tests to classify ancient historical documents based on the physical materials of their writing surfaces. However, these tests, often performed on-site, requires actual access to the manuscript…

Computer Vision and Pattern Recognition · Computer Science 2023-04-13 Thomas Reynolds , Maruf A. Dhali , Lambert Schomaker

The rapidly expanding corpus of medical research literature presents major challenges in the understanding of previous work, the extraction of maximum information from collected data, and the identification of promising research directions.…

Computation and Language · Computer Science 2016-07-19 Victor Andrei , Ognjen Arandjelovic

The GIIDA project aims to develop a digital infrastructure for the spatial information within CNR. It is foreseen to use semantic-oriented technologies to ease information modeling and connecting, according to international standards like…