English
Related papers

Related papers: False Discovery estimation in Record Linkage

200 papers

Concise and meaningful method names are crucial for program comprehension and maintenance. However, method names may become inconsistent with their corresponding implementations, causing confusion and errors. Several deep learning…

Software Engineering · Computer Science 2025-01-23 Taiming Wang , Yuxia Zhang , Lin Jiang , Yi Tang , Guangjie Li , Hui Liu

This paper studies federated learning (FL)--especially cross-silo FL--with data from people who do not trust the server or other silos. In this setting, each silo (e.g. hospital) has data from different people (e.g. patients) and must…

Machine Learning · Computer Science 2024-11-26 Andrew Lowy , Meisam Razaviyayn

Multiple hypothesis testing, a situation when we wish to consider many hypotheses, is a core problem in statistical inference that arises in almost every scientific field. In this setting, controlling the false discovery rate (FDR), which…

Statistics Theory · Mathematics 2019-03-19 Shiyun Chen , Shiva Kasiviswanathan

The false discovery rate (FDR) and the false non-discovery rate (FNR), defined as the expected false discovery proportion (FDP) and the false non-discovery proportion (FNP), are the most popular benchmarks for multiple testing. Despite the…

Statistics Theory · Mathematics 2025-09-03 Yutong Nie , Yihong Wu

While traditional multiple testing procedures prohibit adaptive analysis choices made by users, Goeman and Solari (2011) proposed a simultaneous inference framework that allows users such flexibility while preserving high-probability bounds…

Statistics Theory · Mathematics 2021-01-05 Eugene Katsevich , Aaditya Ramdas

Causal discovery is crucial for causal inference in observational studies, as it can enable the identification of valid adjustment sets (VAS) for unbiased effect estimation. However, global causal discovery is notoriously hard in the…

Machine Learning · Statistics 2024-06-04 Jacqueline Maasch , Weishen Pan , Shantanu Gupta , Volodymyr Kuleshov , Kyra Gan , Fei Wang

There has been recent interest in extending the ideas of False Discovery Rates (FDR) to variable selection in regression settings. Traditionally the FDR in these settings has been defined in terms of the coefficients of the full regression…

Methodology · Statistics 2013-02-12 Max Grazier G'Sell , Trevor Hastie , Robert Tibshirani

False discovery rate (FDR) is commonly used for correction for multiple testing in neuroimaging studies. However, when using two-tailed tests, making directional inferences about the results can lead to a vastly inflated error rate, even…

Methodology · Statistics 2025-12-16 Anderson M. Winkler , Paul A. Taylor , Thomas E. Nichols , Chris Rorden

Unstructured data is pervasive, but analytical queries demand structured representations, creating a significant extraction challenge. Existing methods like RAG lack schema awareness and struggle with cross-document alignment, leading to…

Databases · Computer Science 2025-11-05 Daren Chao , Kaiwen Chen , Naiqing Guan , Nick Koudas

In many applications, the process of identifying a specific feature of interest often involves testing multiple hypotheses for their joint statistical significance. Examples include mediation analysis which simultaneously examines the…

Methodology · Statistics 2023-05-30 Linsui Deng , Kejun He , Xianyang Zhang

As the volume and complexity of data continue to expand across various scientific disciplines, the need for robust methods to account for the multiplicity of comparisons has grown widespread. A popular measure of type 1 error rate in…

Methodology · Statistics 2024-11-19 Jianliang He , Bowen Gang , Luella Fu

Out of the participants in a randomized experiment with anticipated heterogeneous treatment effects, is it possible to identify which subjects have a positive treatment effect? While subgroup analysis has received attention, claims about…

Methodology · Statistics 2024-05-14 Boyan Duan , Larry Wasserman , Aaditya Ramdas

We consider a multi-object detection problem over a sensor network (SNET) with limited range sensors. This problem complements the widely considered decentralized detection problem where all sensors observe the same object. While the…

Information Theory · Computer Science 2016-11-17 Erhan B. Ermis , Venkatesh Saligrama

We propose a ranking and selection procedure to prioritize relevant predictors and control false discovery proportion (FDP) of variable selection. Our procedure utilizes a new ranking method built upon the de-sparsified Lasso estimator. We…

Methodology · Statistics 2018-12-12 X. Jessie Jeng , Xiongzhi Chen

When outcomes are missing for reasons beyond an investigator's control, there are two different ways to adjust a parameter estimate for covariates that may be related both to the outcome and to missingness. One approach is to model the…

Methodology · Statistics 2008-12-18 Joseph D. Y. Kang , Joseph L. Schafer

Privacy-preserving record linkage (PPRL), the problem of identifying records that correspond to the same real-world entity across several data sources held by different parties without revealing any sensitive information about these…

Databases · Computer Science 2016-12-30 Dinusha Vatsalan , Peter Christen

The mitigation of false positives is an important issue when conducting multiple hypothesis testing. The most popular paradigm for false positives mitigation in high-dimensional applications is via the control of the false discovery rate…

Methodology · Statistics 2018-07-17 Hien D. Nguyen , Yohan Yee , Geoffrey J. McLachlan , Jason P. Lerch

The multiple testing procedure plays an important role in detecting the presence of spatial signals for large-scale imaging data. Typically, the spatial signals are sparse but clustered. This paper provides empirical evidence that for a…

Statistics Theory · Mathematics 2011-03-11 Chunming Zhang , Jianqing Fan , Tao Yu

Recent tools for interactive data exploration significantly increase the chance that users make false discoveries. The crux is that these tools implicitly allow the user to test a large body of different hypotheses with just a few clicks…

Databases · Computer Science 2016-12-06 Zheguang Zhao , Lorenzo De Stefani , Emanuel Zgraggen , Carsten Binnig , Eli Upfal , Tim Kraska

Multiple imputation (MI) inference handles missing data by imputing the missing values $m$ times, and then combining the results from the $m$ complete-data analyses. However, the existing method for combining likelihood ratio tests (LRTs)…

Statistics Theory · Mathematics 2022-01-03 Kin Wai Chan , Xiao-Li Meng