中文
相关论文

相关论文: Bayesian Estimation of Bipartite Matchings for Rec…

200 篇论文

Merging datafiles containing information on overlapping sets of entities is a challenging task in the absence of unique identifiers, and is further complicated when some entities are duplicated in the datafiles. Most approaches to this…

统计方法学 · 统计学 2021-10-11 Serge Aleshin-Guendel , Mauricio Sadinle

In many settings, a data curator links records from two files to produce datasets that are shared with secondary analysts. Analysts use the linked files to estimate models of interest, such as regressions. Such two-stage approaches do not…

统计方法学 · 统计学 2025-11-18 Xueyan Hu , Jerome P. Reiter

Researchers are often interested in linking individuals between two datasets that lack a common unique identifier. Matching procedures often struggle to match records with common names, birthplaces or other field values. Computational…

统计方法学 · 统计学 2021-06-14 Thomas Stringham

Probabilistic record linkage is often used to match records from two files, in particular when the variables common to both files comprise imperfectly measured identifiers like names and demographic variables. We consider bipartite record…

统计方法学 · 统计学 2023-12-06 Eric A. Bai , Olivier Binette , Jerome P. Reiter

In many applications, researchers seek to identify overlapping entities across multiple data files. Record linkage algorithms facilitate this task, in the absence of unique identifiers. As these algorithms rely on semi-identifying…

统计方法学 · 统计学 2026-04-24 Gauri Kamat , Roee Gutman

Multiple-systems or capture-recapture estimation are common techniques for population size estimation, particularly in the quantitative study of human rights violations. These methods rely on multiple samples from the population, along with…

统计方法学 · 统计学 2018-12-27 Mauricio Sadinle

Record linkage (de-duplication or entity resolution) is the process of merging noisy databases to remove duplicate entities. While record linkage removes duplicate entities from such databases, the downstream task is any inferential,…

统计方法学 · 统计学 2018-10-12 Rebecca C. Steorts , Andrea Tancredi , Brunero Liseo

In many scenarios, the observational data needed for causal inferences are spread over two data files. In particular, we consider scenarios where one file includes covariates and the treatment measured on one set of individuals, and a…

统计方法学 · 统计学 2020-09-22 Sharmistha Guha , Jerome P. Reiter , Andrea Mercatanti

We propose an unsupervised approach for linking records across arbitrarily many files, while simultaneously detecting duplicate records within files. Our key innovation involves the representation of the pattern of links between records as…

统计方法学 · 统计学 2015-11-03 Rebecca C. Steorts , Rob Hall , Stephen E. Fienberg

In many healthcare and social science applications, information about units is dispersed across multiple data files. Linking records across files is necessary to estimate the associations of interest. Common record linkage algorithms only…

统计方法学 · 统计学 2024-06-25 Gauri Kamat , Mingyang Shan , Roee Gutman

We propose and illustrate a hierarchical Bayesian approach for matching statistical records observed on different occasions. We show how this model can be profitably adopted both in record linkage problems and in capture--recapture setups,…

应用统计 · 统计学 2011-07-29 Andrea Tancredi , Brunero Liseo

Existing file linkage methods may produce sub-optimal results because they consider neither the interactions between different pairs of matched records nor relationships between variables that are exclusive to one of the files. In addition,…

统计计算 · 统计学 2021-09-28 Edwin Farley , Roee Gutman

Probabilistic record linkage (PRL) is the process of determining which records in two databases correspond to the same underlying entity in the absence of a unique identifier. Bayesian solutions to this problem provide a powerful mechanism…

统计方法学 · 统计学 2017-12-05 Brendan S. McVeigh , Jared S. Murray

We propose a novel unsupervised approach for linking records across arbitrarily many files, while simultaneously detecting duplicate records within files. Our key innovation is to represent the pattern of links between records as a {\em…

统计计算 · 统计学 2014-03-04 Rebecca C. Steorts , Rob Hall , Stephen E. Fienberg

Finding duplicates in homicide registries is an important step in keeping an accurate account of lethal violence. This task is not trivial when unique identifiers of the individuals are not available, and it is especially challenging when…

应用统计 · 统计学 2015-02-04 Mauricio Sadinle

Databases often contain corrupted, degraded, and noisy data with duplicate entries across and within each database. Such problems arise in citations, medical databases, genetics, human rights databases, and a variety of other applied…

统计方法学 · 统计学 2015-04-29 Rebecca C. Steorts

In record linkage (RL), or exact file matching, the goal is to identify the links between entities with information on two or more files. RL is an important activity in areas including counting the population, enhancing survey frames and…

统计理论 · 数学 2012-12-21 Michael D. Larsen

The increased prevalence of observational data and the need to integrate information from multiple sources are critical challenges in contemporary data analysis. Record linkage is a widely used tool for combining datasets in the absence of…

统计方法学 · 统计学 2025-12-17 Martin Slawski

In theory, the probabilistic linkage method provides two distinct advantages over non-probabilistic methods, including minimal rates of linkage error and accurate measures of these rates for data users. However, implementations can fall…

统计方法学 · 统计学 2019-11-06 Abel Dasylva , Arthur Goussanou , David Ajavon , Hanan Abousaleh

Entity resolution (record linkage or deduplication) is the process of identifying and linking duplicate records in databases. In this paper, we propose a Bayesian graphical approach for entity resolution that links records to latent…

统计方法学 · 统计学 2023-01-10 Neil G. Marchant , Benjamin I. P. Rubinstein , Rebecca C. Steorts
‹ 上一页 1 2 3 10 下一页 ›