English
Related papers

Related papers: Variable Selection for Multi-Source Count Data wit…

200 papers

Power and reproducibility are key to enabling refined scientific discoveries in contemporary big data applications with general high-dimensional nonlinear models. In this paper, we provide theoretical foundations on the power and robustness…

Statistics Theory · Mathematics 2017-09-04 Yingying Fan , Emre Demirkaya , Gaorong Li , Jinchi Lv

Genome-wide association studies (GWAS) often find association signals between many genetic variants and traits of interest in a genomic region. Functional annotations of these variants provide valuable prior information that helps…

Methodology · Statistics 2026-01-07 Xiangyu Zhang , Lijun Wang , Changjun Li , Chen Lin , Hongyu Zhao

We extend the knockoffs method for selecting predictors to clustered data (cross-sectional or repeated measures). In the setting of clustered data, variable selection is complex because some predictors are measured at the observation level…

Methodology · Statistics 2026-02-24 Silvia Bacci , Leonardo Grilli , Carla Rampichini

Understanding the association between dietary patterns and health outcomes, such as the cancer risk, is crucial to inform public health guidelines and shaping future dietary interventions. However, dietary intake data present several…

Methodology · Statistics 2025-10-10 Blake Hansen , Dafne Zorzetto , Valeria Edefonti , Roberta De Vito

We consider the problem of assessing the importance of multiple variables or factors from a dataset when side information is available. In principle, using side information can allow the statistician to pay attention to variables with a…

Methodology · Statistics 2020-01-23 Zhimei Ren , Emmanuel Candès

High-dimensional variable selection has emerged as one of the prevailing statistical challenges in the big data revolution. Many variable selection methods have been adapted for identifying single nucleotide polymorphisms (SNPs) linked to…

Methodology · Statistics 2024-08-21 Justin J. Van Ee , Diana Gamba , Jesse R. Lasky , Megan L. Vahsen , Mevin B. Hooten

In many research fields, researchers aim to identify significant associations between a set of explanatory variables and a response while controlling the FDR. The Knockoff filter has been recently proposed in the frequentist paradigm to…

Methodology · Statistics 2026-04-22 Lorenzo Focardi-Olmi , Anna Gottard , Michele Guindani , Marina Vannucci

We propose a new method to learn the structure of a Gaussian graphical model with finite sample false discovery rate control. Our method builds on the knockoff framework of Barber and Cand\`{e}s for linear models. We extend their approach…

Methodology · Statistics 2021-04-20 Jinzhou Li , Marloes H. Maathuis

Multiple comparisons in hypothesis testing often encounter structural constraints in various applications. For instance, in structural Magnetic Resonance Imaging for Alzheimer's Disease, the focus extends beyond examining atrophic brain…

Methodology · Statistics 2023-11-08 Yang Cao , Xinwei Sun , Yuan Yao

Continuous improvement in medical imaging techniques allows the acquisition of higher-resolution images. When these are used in a predictive setting, a greater number of explanatory variables are potentially related to the dependent…

Statistics Theory · Mathematics 2019-03-13 Tuan-Binh Nguyen , Jérôme-Alexis Chevalier , Bertrand Thirion

We address challenges in variable selection with highly correlated data that are frequently present in finance, economics, but also in complex natural systems as e.g. weather. We develop a robustified version of the knockoff framework,…

Econometrics · Economics 2022-06-14 Konstantin Görgen , Abdolreza Nazemi , Melanie Schienle

Many clinical endpoint measures, such as the number of standard drinks consumed per week or the number of days that patients stayed in the hospital, are count data with excessive zeros. However, the zero-inflated nature of such outcomes is…

Applications · Statistics 2022-07-14 Zhengyang Zhou , Minge Xie , David Huh , Eun-Young Mun

We propose the group knockoff filter, a method for false discovery rate control in a linear regression setting where the features are grouped, and we would like to select a set of relevant groups which have a nonzero effect on the response.…

Methodology · Statistics 2016-02-12 Ran Dai , Rina Foygel Barber

Zero-inflated datasets, which have an excess of zero outputs, are commonly encountered in problems such as climate or rare event modelling. Conventional machine learning approaches tend to overestimate the non-zeros leading to poor…

Machine Learning · Statistics 2018-03-15 Pashupati Hegde , Markus Heinonen , Samuel Kaski

This paper proposes a model-free and data-adaptive feature screening method for ultra-high dimensional datasets. The proposed method is based on the projection correlation which measures the dependence between two random vectors. This…

Methodology · Statistics 2021-02-16 Wanjun Liu , Yuan Ke , Jingyuan Liu , Runze Li

High-dimensional sparse matrix data frequently arise in various applications. A notable example is the weighted word-word co-occurrence count data, which summarizes the weighted frequency of word pairs appearing within the same context…

Machine Learning · Computer Science 2025-01-03 Taejoon Kim , Haiyan Wang

Standard MCMC methods can scale poorly to big data settings due to the need to evaluate the likelihood at each iteration. There have been a number of approximate MCMC algorithms that use sub-sampling ideas to reduce this computational…

Computation · Statistics 2020-09-29 Joris Bierkens , Paul Fearnhead , Gareth Roberts

The false discovery rate (FDR)---the expected fraction of spurious discoveries among all the discoveries---provides a popular statistical assessment of the reproducibility of scientific studies in various disciplines. In this work, we…

Machine Learning · Statistics 2015-11-10 Weijie Su , Junyang Qian , Linxi Liu

Addressing the simultaneous identification of contributory variables while controlling the false discovery rate (FDR) in high-dimensional data is a crucial statistical challenge. In this paper, we propose a novel model-free variable…

Methodology · Statistics 2024-04-23 Yixin Han , Xu Guo , Changliang Zou

We propose a unified probabilistic framework for sparse count tensors with excess zeros, motivated by single-cell Hi-C data. The observed data are naturally represented as a three-way tensor indexed by genomic loci pairs and cells,…

Methodology · Statistics 2026-04-27 Elena Tuzhilina , Yaoming Zhen