English
Related papers

Related papers: Factor-Adjusted Multiple Testing for High-Dimensio…

200 papers

Large language models (LLMs) show amazing performance on many domain-specific tasks after fine-tuning with some appropriate data. However, many domain-specific data are privately distributed across multiple owners. Thus, this dilemma raises…

Machine Learning · Computer Science 2024-06-26 Feijie Wu , Zitao Li , Yaliang Li , Bolin Ding , Jing Gao

We consider mediated effects of an exposure, X on an outcome, Y, via a mediator, M, under no unmeasured confounding assumptions in the setting where models for the conditional expectation of the mediator and outcome are partially linear. We…

Methodology · Statistics 2025-01-08 Oliver Hines , Stijn Vansteelandt , Karla Diaz-Ordaz

In mediation analysis, the exposure often influences the mediating effect, i.e., there is an interaction between exposure and mediator on the dependent variable. When the mediator is high-dimensional, it is necessary to identify non-zero…

Methodology · Statistics 2023-08-03 Ruiyang Li , Xi Zhu , Seonjoo Lee

Much effort has been done to control the "false discovery rate" (FDR) when $m$ hypotheses are tested simultaneously. The FDR is the expectation of the "false discovery proportion" $\text{FDP}=V/R$ given by the ratio of the number of false…

Statistics Theory · Mathematics 2018-01-09 Marc Ditzhaus , Arnold Janssen

Adaptive interventions, aka dynamic treatment regimens, are sequences of pre-specified decision rules that guide the provision of treatment for an individual given information about their baseline and evolving needs, including in response…

Methodology · Statistics 2024-05-02 Wenchu Pan , Daniel Almirall , Amy M. Kilbourne , Andrew Quanbeck , Lu Wang

Federated learning (FL) with noisy labels poses a significant challenge. Existing methods designed for handling noisy labels in centralized learning tend to lose their effectiveness in the FL setting, mainly due to the small dataset size…

Machine Learning · Computer Science 2024-01-11 Lei Wang , Jieming Bian , Jie Xu

Large-scale hypothesis testing has become a ubiquitous problem in high-dimensional statistical inference, with broad applications in various scienfitic disciplines. One relevant application is constituted by imaging mass spectrometry (IMS)…

Methodology · Statistics 2021-08-19 Vladimir Vutov , Thorsten Dickhaus

Privacy-preserving model co-training in medical research is often hindered by server-dependent architectures incompatible with protected hospital data systems and by the predominant focus on relative effect measures (hazard ratios) which…

Machine Learning · Statistics 2026-01-22 Ziwen Wang , Siqi Li , Marcus Eng Hock Ong , Nan Liu

The mitigation of false positives is an important issue when conducting multiple hypothesis testing. The most popular paradigm for false positives mitigation in high-dimensional applications is via the control of the false discovery rate…

Methodology · Statistics 2018-07-17 Hien D. Nguyen , Yohan Yee , Geoffrey J. McLachlan , Jason P. Lerch

We propose a novel multiple testing methodology for controlling the false discovery rate (FDR) in high-dimensional linear models that integrates model-X knockoff techniques with debiased penalized regression estimators. At the foundation of…

Methodology · Statistics 2026-03-17 Jinyuan Chang , Chenlong Li , Cheng Yong Tang , Zhengtian Zhu

Mediation analysis aims to decipher the underlying causal mechanisms between an exposure, an outcome, and intermediate variables called mediators. Initially developed for fixed-time mediator and outcome, it has been extended to the…

Methodology · Statistics 2025-01-15 K. Le Bourdonnec , L. Valeri , C. Proust-Lima

We address the multiple testing problem under the assumption that the true/false hypotheses are driven by a Hidden Markov Model (HMM), which is recognized as a fundamental setting to model multiple testing under dependence since the seminal…

Methodology · Statistics 2021-05-04 Marie Perrot-Dockès , Gilles Blanchard , Pierre Neuvial , Etienne Roquain

As Large Language Models (LLMs) transition from static tools to autonomous agents, traditional evaluation benchmarks that measure performance on downstream tasks are becoming insufficient. These methods fail to capture the emergent social…

Artificial Intelligence · Computer Science 2025-10-03 Zarreen Reza

Non-independent and identically distributed (non-IID) data is a key challenge in federated learning (FL), which usually hampers the optimization convergence and the performance of FL. Existing data augmentation methods based on federated…

Machine Learning · Computer Science 2023-01-13 Shaoming Duan , Chuanyi Liu , Peiyi Han , Tianyu He , Yifeng Xu , Qiyuan Deng

Selecting relevant features associated with a given response variable is an important issue in many scientific fields. Quantifying quality and uncertainty of a selection result via false discovery rate (FDR) control has been of recent…

Methodology · Statistics 2020-12-17 Chenguang Dai , Buyu Lin , Xin Xing , Jun S. Liu

Large Language Models (LLMs) excel at human-like language generation but often embed and amplify implicit, intersectional biases, especially under persona-driven contexts. Existing bias audits rely on static, embedding-based tests (CEAT,…

Computation and Language · Computer Science 2026-04-09 Nandini Arimanda , Achyuth Mukund , Sakthi Balan Muthiah , Rajesh Sharma

This paper develops a general framework for controlling the false discovery rate (FDR) in multiple testing of Gaussian means against two-sided alternatives. The widely used Benjamini-Hochberg (BH) procedure provides exact FDR control under…

Methodology · Statistics 2025-11-26 Deepra Ghosh , Sanat K. Sarkar

We develop a flexible feature selection framework based on deep neural networks that approximately controls the false discovery rate (FDR), a measure of Type-I error. The method applies to architectures whose first layer is fully connected.…

Machine Learning · Statistics 2026-02-10 Kazuma Sawaya

Parameter-efficient fine-tuning (PEFT) methods typically assume that Large Language Models (LLMs) are trained on data from a single device or client. However, real-world scenarios often require fine-tuning these models on private data…

Machine Learning · Computer Science 2025-06-03 Sajjad Ghiasvand , Yifan Yang , Zhiyu Xue , Mahnoosh Alizadeh , Zheng Zhang , Ramtin Pedarsani

This paper develops a semiparametric Bayesian instrumental variable analysis method for estimating the causal effect of an endogenous variable when dealing with unobserved confounders and measurement errors with partly interval-censored…

Methodology · Statistics 2025-01-28 Elvis Han Cui , Xuyang Lu , Jin Zhou , Hua Zhou , Gang Li