English
Related papers

Related papers: Large Scale Genealogical Information Extraction Fr…

200 papers

This paper describes a series of automated data validation tests for datasets detailing charity financial information, political donations, and government lobbying in Canada. We motivate and document a series of 200 tests that check the…

Methodology · Statistics 2023-09-25 Lindsay Katz , Callandra Moore

Large Language Models (LLMs) perform outstandingly in various downstream tasks, and the use of the Retrieval-Augmented Generation (RAG) architecture has been shown to improve performance for legal question answering (Nuruzzaman and Hussain,…

Computation and Language · Computer Science 2024-10-15 David Beauchemin , Zachary Gagnon , Ricahrd Khoury

When extracting information from handwritten documents, text transcription and named entity recognition are usually faced as separate subsequent tasks. This has the disadvantage that errors in the first module affect heavily the performance…

Computer Vision and Pattern Recognition · Computer Science 2018-03-23 Manuel Carbonell , Mauricio Villegas , Alicia Fornés , Josep Lladós

Computational reproducibility is central to scientific credibility, yet verifying published results at scale remains costly. We develop an AI-assisted workflow for automated full-paper replication -- retrieving materials, reconstructing…

Econometrics · Economics 2026-03-27 Yiqing Xu , Leo Yang Yang

The digital age allows data collection to be done on a large scale and at low cost. This is the case of genealogy trees, which flourish on numerous digital platforms thanks to the collaboration of a mass of individuals wishing to trace…

Physics and Society · Physics 2018-07-25 Arthur Charpentier , Ewen Gallic

Driven by the popularity of television shows such as Who Do You Think You Are? many millions of users have uploaded their family tree to web projects such as WikiTree. Analysis of this corpus enables us to investigate genealogy…

Social and Information Networks · Computer Science 2014-09-02 Michael Fire , Thomas Chesney , Yuval Elovici

In this paper, we investigate the use of Convolutional Neural Networks for counting the number of records in historical handwritten documents. With this work we demonstrate that training the networks only with synthetic images allows us to…

Computer Vision and Pattern Recognition · Computer Science 2017-11-21 Samuele Capobianco , Simone Marinai

We describe the project REE-HDSC and outline our efforts to improve the quality of named entities extracted automatically from texts generated by hand-written text recognition (HTR) software. We describe a six-step processing pipeline and…

Computation and Language · Computer Science 2024-04-08 Erik Tjong Kim Sang

The context of this paper is the creation of large uniform archaeological datasets from heterogeneous published resources, such as find catalogues - with the help of AI and Big Data. The paper is concerned with the challenge of consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Kevin Klein , Antoine Muller , Alyssa Wohde , Alexander V. Gorelik , Volker Heyd , Ralf Lämmel , Yoan Diekmann , Maxime Brami

Computer-based assessments routinely generate detailed interaction logs -- commonly referred to as process data -- that record every action a respondent performs during task completion, yet systematic preprocessing guidance, integrated…

Applications · Statistics 2026-04-21 Daeun Hwangbo , Junyeong Park , Minjeong Jeon , Ick Hoon Jin

Query Autocomplete (QAC) is a critical feature in modern search engines, facilitating user interaction by predicting search queries based on input prefixes. Despite its widespread adoption, the absence of large-scale, realistic datasets has…

Information Retrieval · Computer Science 2024-11-08 Dante Everaert , Rohit Patki , Tianqi Zheng , Christopher Potts

Handwritten text recognition has been widely studied in the last decades for its numerous applications. Nowadays, the state-of-the-art approach consists in a three-step process. The document is segmented into text lines, which are then…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Denis Coquenet

Genealogical networks, also known as family trees or population pedigrees, are commonly studied by genealogists wanting to know about their ancestry, but they also provide a valuable resource for disciplines such as digital demography,…

Social and Information Networks · Computer Science 2018-02-19 Eric Malmi , Aristides Gionis , Arno Solin

Automatic extraction of procedural graphs from documents creates a low-cost way for users to easily understand a complex procedure by skimming visual graphs. Despite the progress in recent studies, it remains unanswered: whether the…

Computation and Language · Computer Science 2024-08-09 Weihong Du , Wenrui Liao , Hongru Liang , Wenqiang Lei

When knowledge graphs (KGs) are automatically extracted from text, are they accurate enough for downstream analysis? Unfortunately, current annotated datasets can not be used to evaluate this question, since their KGs are highly…

Computation and Language · Computer Science 2025-05-19 Erica Cai , Sean McQuade , Kevin Young , Brendan O'Connor

The detection of molecular signatures of selection is one of the major concerns of modern population genetics. A widely used strategy in this context is to compare samples from several populations, and to look for genomic regions with…

Populations and Evolution · Quantitative Biology 2013-01-24 Marìa Inès Fariello , Simon Boitard , Hugo Naya , Magali SanCristobal , Bertrand Servin

Researchers are often interested in linking individuals between two datasets that lack a common unique identifier. Matching procedures often struggle to match records with common names, birthplaces or other field values. Computational…

Methodology · Statistics 2021-06-14 Thomas Stringham

This article presents and validates an ideal, four-stage workflow for the high-accuracy transcription and analysis of challenging medieval legal documents. The process begins with a specialized Handwritten Text Recognition (HTR) model,…

Digital Libraries · Computer Science 2025-07-08 Joshua D. Isom

In this paper, we explore different ways of training a model for handwritten text recognition when multiple imperfect or noisy transcriptions are available. We consider various training configurations, such as selecting a single…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Solène Tarride , Tristan Faine , Mélodie Boillet , Harold Mouchère , Christopher Kermorvant

In this report, we present our findings from benchmarking experiments for information extraction on historical handwritten marriage records Esposalles from IEHHR - ICDAR 2017 robust reading competition. The information extraction is modeled…

Computer Vision and Pattern Recognition · Computer Science 2018-07-18 Animesh Prasad , Hervé Déjean , Jean-Luc Meunier , Max Weidemann , Johannes Michael , Gundram Leifert