English
Related papers

Related papers: False Discovery estimation in Record Linkage

200 papers

In a context of multiple hypothesis testing, we provide several new exact calculations related to the false discovery proportion (FDP) of step-up and step-down procedures. For step-up procedures, we show that the number of erroneous…

Statistics Theory · Mathematics 2011-06-29 Etienne Roquain , Fanny Villers

The sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL). A promising approach to addressing this challenge is distributionally robust RL,…

Machine Learning · Computer Science 2024-11-05 Miao Lu , Han Zhong , Tong Zhang , Jose Blanchet

Intuitively, an ideal collaborative filtering (CF) model should learn from users' full rankings over all items to make optimal top-K recommendations. Due to the absence of such full rankings in practice, most CF models rely on pairwise loss…

Information Retrieval · Computer Science 2024-12-25 Yuhan Zhao , Rui Chen , Li Chen , Shuang Zhang , Qilong Han , Hongtao Song

Modern applications of conformal inference to multiple testing problems, such as outlier detection and candidate selection, often involve selecting test samples whose conformal p-values fall below a threshold. The quality of such methods is…

Methodology · Statistics 2026-05-21 Ziang Song , Ying Jin , Emmanuel J. Candès

Probabilistic record linkage is often used to match records from two files, in particular when the variables common to both files comprise imperfectly measured identifiers like names and demographic variables. We consider bipartite record…

Methodology · Statistics 2023-12-06 Eric A. Bai , Olivier Binette , Jerome P. Reiter

Probabilistic record linkage (PRL) is the process of determining which records in two databases correspond to the same underlying entity in the absence of a unique identifier. Bayesian solutions to this problem provide a powerful mechanism…

Methodology · Statistics 2017-12-05 Brendan S. McVeigh , Jared S. Murray

There are a number of ways to test for the absence/presence of a spatial signal in a completely observed fine-resolution image. One of these is a powerful nonparametric procedure called Enhanced False Discovery Rate (EFDR). A drawback of…

Methodology · Statistics 2020-10-20 Hsin-Cheng Huang , Noel Cressie , Andrew Zammit-Mangion , Guowen Huang

Simultaneously performing variable selection and inference in high-dimensional regression models is an open challenge in statistics and machine learning. The increasing availability of vast amounts of variables requires the adoption of…

Methodology · Statistics 2025-05-08 Marco Molinari , Magne Thoresen

Matching households and individuals across different databases poses challenges due to the lack of unique identifiers, typographical errors, and changes in attributes over time. Record linkage tools play a crucial role in overcoming these…

Applications · Statistics 2024-04-09 Thais Pacheco Menezes , Thomas Brendan Murphy , Michael Fop

Deep learning-based linkage of records across different databases is becoming increasingly useful in data integration and mining applications to discover new insights from multiple sources of data. However, due to privacy and…

Cryptography and Security · Computer Science 2022-11-07 Thilina Ranbaduge , Dinusha Vatsalan , Ming Ding

The probability of false discovery proportion (FDP) exceeding $\gamma\in[0,1)$, defined as $\gamma$-FDP, has received much attention as a measure of false discoveries in multiple testing. Although this measure has received acceptance due to…

Statistics Theory · Mathematics 2014-06-03 Wenge Guo , Li He , Sanat K. Sarkar

Local differential privacy (LDP) has emerged as a promising paradigm for privacy-preserving data collection in distributed systems, where users contribute multi-dimensional records with potentially correlated attributes. Recent work has…

Cryptography and Security · Computer Science 2025-08-20 Sandaru Jayawardana , Sennur Ulukus , Ming Ding , Kanchana Thilakarathna

We develop a new class of distribution--free multiple testing rules for false discovery rate (FDR) control under general dependence. A key element in our proposal is a symmetrized data aggregation (SDA) approach to incorporating the…

Methodology · Statistics 2021-05-27 Lilun Du , Xu Guo , Wenguang Sun , Changliang Zou

This paper extends the theory of false discovery rates (FDR) pioneered by Benjamini and Hochberg [J. Roy. Statist. Soc. Ser. B 57 (1995) 289-300]. We develop a framework in which the False Discovery Proportion (FDP)--the number of false…

Statistics Theory · Mathematics 2007-06-13 Christopher Genovese , Larry Wasserman

The false discovery rate (FDR)---the expected fraction of spurious discoveries among all the discoveries---provides a popular statistical assessment of the reproducibility of scientific studies in various disciplines. In this work, we…

Machine Learning · Statistics 2015-11-10 Weijie Su , Junyang Qian , Linxi Liu

In the community of Linked Data, anyone can publish their data as Linked Data on the web because of the openness of the Semantic Web. As such, RDF (Resource Description Framework) triples described the same real-world entity can be obtained…

Databases · Computer Science 2017-04-25 Wenqiang Liu

In theory, the probabilistic linkage method provides two distinct advantages over non-probabilistic methods, including minimal rates of linkage error and accurate measures of these rates for data users. However, implementations can fall…

Methodology · Statistics 2019-11-06 Abel Dasylva , Arthur Goussanou , David Ajavon , Hanan Abousaleh

Record linkage is aimed at the accurate and efficient identification of records that represent the same entity within or across disparate databases. It is a fundamental task in data integration and increasingly required for accurate…

Databases · Computer Science 2021-04-21 Thilina Ranbaduge , Peter Christen , Rainer Schnell

We provide an approach to exploratory data analysis in matched observational studies with a single intervention and multiple endpoints. In such settings, the researcher would like to explore evidence for actual treatment effects among these…

Methodology · Statistics 2025-12-10 Mengqi Lin , Colin Fogarty

Existing differentially private (DP) synthetic data generation mechanisms typically assume a single-source table. In practice, data is often distributed across multiple tables with relationships across tables. In this paper, we introduce…

Machine Learning · Computer Science 2025-01-22 Kaveh Alimohammadi , Hao Wang , Ojas Gulati , Akash Srivastava , Navid Azizan