English
Related papers

Related papers: Factor-Adjusted Multiple Testing for High-Dimensio…

200 papers

Federated learning has attracted significant recent attention due to its applicability across a wide range of settings where data is collected and analyzed across disparate locations. In this paper, we study federated nonparametric…

Statistics Theory · Mathematics 2024-06-12 T. Tony Cai , Abhinav Chakraborty , Lasse Vuursteen

Improving the accessibility of psychotherapy with the aid of Large Language Models (LLMs) is garnering a significant attention in recent years. Recognizing cognitive distortions from the interviewee's utterances can be an essential part of…

Computation and Language · Computer Science 2024-03-22 Sehee Lim , Yejin Kim , Chi-Hyun Choi , Jy-yong Sohn , Byung-Hoon Kim

In recent years, multiple hypothesis testing has come to the forefront of statistical research, ostensibly in relation to applications in genomics and some other emerging fields. The false discovery rate (FDR) and its variants provide very…

Statistics Theory · Mathematics 2008-12-18 Subhashis Ghosal , Anindya Roy , Yongqiang Tang

Federated learning (FL) is a promising framework for learning from distributed data while maintaining privacy. The development of efficient FL algorithms encounters various challenges, including heterogeneous data and systems, limited…

Machine Learning · Computer Science 2024-08-01 Yongcun Song , Ziqi Wang , Enrique Zuazua

This paper addresses the problem of frequency-domain inter-carrier interference (ICI) mitigation for differential orthogonal frequency-division multiplexing (OFDM) systems. The classical fractional fast Fourier transform (F-FFT), adopting…

Information Theory · Computer Science 2021-10-12 Jihui Qiu , Yuzhou Li , Yunlong Huang , Yimeng Wang , Lingyu Gu

False discovery rate (FDR) control is a popular approach for maintaining the integrity of statistical analyses, especially in high-dimensional data settings, where multiple comparisons increase the risk of false positives. FDR control has…

Signal Processing · Electrical Eng. & Systems 2026-03-03 Fabian Scheidt , Jasin Machkour , Michael Muma

Controlled Direct Effect (CDE) is one of the causal estimands used to evaluate both exposure and mediation effects on an outcome. When there are unmeasured confounders existing between the mediator and the outcome, the ordinary…

Methodology · Statistics 2024-10-30 Shunichiro Orihara , Shinpei Imori , Kosuke Morikawa , Atsushi Goto , Masataka Taguri

This article considers the problem of multiple hypothesis testing using $t$-tests. The observed data are assumed to be independently generated conditional on an underlying and unknown two-state hidden model. We propose an asymptotically…

Statistics Theory · Mathematics 2011-02-22 Hongyuan Cao , Michael R. Kosorok

Federated learning (FL) enables multiple clients to collaboratively train deep learning models while considering sensitive local datasets' privacy. However, adversaries can manipulate datasets and upload models by injecting triggers for…

Machine Learning · Computer Science 2023-07-04 Zekai Chen , Fuyi Wang , Zhiwei Zheng , Ximeng Liu , Yujie Lin

High-dimensional sparse generalized linear models (GLMs) have emerged in the setting that the number of samples and the dimension of variables are large, and even the dimension of variables grows faster than the number of samples. False…

Statistics Theory · Mathematics 2021-05-04 Chang Cui , Jinzhu Jia , Yijun Xiao , Huiming Zhang

Given a nonparametric Hidden Markov Model (HMM) with two states, the question of constructing efficient multiple testing procedures is considered, treating one of the states as an unknown null hypothesis. A procedure is introduced, based on…

Statistics Theory · Mathematics 2021-01-12 Kweku Abraham , Ismael Castillo , Elisabeth Gassiat

Understanding the pathways whereby an intervention has an effect on an outcome is a common scientific goal. A rich body of literature provides various decompositions of the total intervention effect into pathway specific effects.…

Methodology · Statistics 2020-01-20 David Benkeser

In this paper we propose a heterogeneous modeling framework which achieves individual-wise feature selection and individualized covariates' effects subgrouping simultaneously. In contrast to conventional model selection approaches, the new…

Methodology · Statistics 2019-06-11 Xiwei Tang , Fei Xue , Annie Qu

In hypothesis testing, a false discovery occurs when a hypothesis is incorrectly rejected due to noise in the sample. When adaptively testing multiple hypotheses, the probability of a false discovery increases as more tests are performed.…

Machine Learning · Statistics 2020-10-22 Wanrong Zhang , Gautam Kamath , Rachel Cummings

Depression is a highly prevalent and disabling condition that incurs substantial personal and societal costs. Current depression diagnosis involves determining the depression severity of a person through self-reported questionnaires or…

Computation and Language · Computer Science 2025-03-27 Aishik Mandal , Dana Atzil-Slonim , Thamar Solorio , Iryna Gurevych

Fact-checking research has extensively explored verification but less so the generation of natural-language explanations, crucial for user trust. While Large Language Models (LLMs) excel in text generation, their capability for producing…

Computation and Language · Computer Science 2024-02-13 Kyungha Kim , Sangyun Lee , Kung-Hsiang Huang , Hou Pong Chan , Manling Li , Heng Ji

MaxT is a highly popular resampling-based multiple testing procedure, which controls the Familywise Error Rate (FWER) and is powerful under dependence. This paper generalizes maxT to what we term ``multi-resolution'' False Discovery…

Methodology · Statistics 2026-05-05 Jesse Hemerik

This paper studies the high-dimensional mixed linear regression (MLR) where the output variable comes from one of the two linear regression models with an unknown mixing proportion and an unknown covariance structure of the random…

Methodology · Statistics 2020-11-10 Linjun Zhang , Rong Ma , T. Tony Cai , Hongzhe Li

Despite demonstrating superior performance across a variety of linguistic tasks, pre-trained large language models (LMs) often require fine-tuning on specific datasets to effectively address different downstream tasks. However, fine-tuning…

Computation and Language · Computer Science 2024-10-02 Zhidong Gao , Yu Zhang , Zhenxiao Zhang , Yanmin Gong , Yuanxiong Guo

Traditional federated learning uses the number of samples to calculate the weights of each client model and uses this fixed weight value to fusion the global model. However, in practical scenarios, each client's device and data…

Machine Learning · Computer Science 2024-03-20 Leiming Chen , Weishan Zhang , Cihao Dong , Sibo Qiao , Ziling Huang , Yuming Nie , Zhaoxiang Hou , Chee Wei Tan