English
Related papers

Related papers: Variable Selection for Multi-Source Count Data wit…

200 papers

This paper proposes a new generalized linear model with the fractional binomial distribution. Zero-inflated Poisson/negative binomial distributions are used for count data with many zeros. To analyze the association of such a count variable…

Methodology · Statistics 2025-08-01 Jeonghwa Lee , Chloe Breece

We introduce DiffKnock, a diffusion-based knockoff framework for high-dimensional feature selection with finite-sample false discovery rate (FDR) control. DiffKnock addresses two key limitations of existing knockoff methods: preserving…

Methodology · Statistics 2025-10-03 Heng Ge , Qing Lu

Multi-omic datasets offer opportunities for improved biomarker discovery in cancer research, but their high dimensionality and limited sample sizes make identifying compact and effective biomarker panels challenging. Feature selection in…

Genomics · Quantitative Biology 2026-04-02 Luca Cattelani , Vittorio Fortino

An important problem in machine learning and statistics is to identify features that causally affect the outcome. This is often impossible to do from purely observational data, and a natural relaxation is to identify features that are…

Machine Learning · Statistics 2019-05-30 Jaime Roquero Gimenez , Amirata Ghorbani , James Zou

The Poisson distribution is often used as a standard model for count data. Quite often, however, such data sets are not well fit by a Poisson model because they have more zeros than are compatible with this model. For these situations, a…

Statistics Theory · Mathematics 2008-12-18 M. J. Bayarri , James O. Berger , Gauri S. Datta

Testing multiple hypotheses of conditional independence with provable error rate control is a fundamental problem with various applications. To infer conditional independence with family-wise error rate (FWER) control when only summary…

Methodology · Statistics 2023-10-17 Catherine Xinrui Yu , Jiaqi Gu , Zhaomeng Chen , Zihuai He

We introduce a new dynamical system for sequentially observed multivariate count data. This model is based on the gamma--Poisson construction---a natural choice for count data---and relies on a novel Bayesian nonparametric prior that ties…

Machine Learning · Statistics 2017-01-23 Aaron Schein , Mingyuan Zhou , Hanna Wallach

Human microbiome studies based on genetic sequencing techniques produce compositional longitudinal data of the relative abundances of microbial taxa over time, allowing to understand, through mixed-effects modeling, how microbial…

Methodology · Statistics 2025-12-23 John Barrera , Cristian Meza , Ana Arribas-Gil

Over the past decades, statisticians and machine-learning researchers have developed literally thousands of new tools for the reduction of high-dimensional data in order to identify the variables most responsible for a particular trait.…

Machine Learning · Statistics 2012-05-31 Chamont Wang , Jana Gevertz , Chaur-Chin Chen , Leonardo Auslender

Identifying which variables do influence a response while controlling false positives pervades statistics and data science. In this paper, we consider a scenario in which we only have access to summary statistics, such as the values of…

Feature selection is critical in machine learning to reduce dimensionality and improve model accuracy and efficiency. The exponential growth in feature space dimensionality for modern datasets directly results in ambiguous samples and…

Quantum Physics · Physics 2023-11-30 Haiyan Wang

Microorganisms play critical roles in human health and disease. It is well known that microbes live in diverse communities in which they interact synergistically or antagonistically. Thus for estimating microbial associations with clinical…

Modern cancer genomics datasets involve widely varying sizes and scales, measurement variables, and correlation structures. A fundamental analytical goal in these high-throughput studies is the development of general statistical techniques…

Methodology · Statistics 2022-04-12 Chiyu Gu , Veerabhadran Baladandayuthapani , Subharup Guha

In many scientific problems, researchers try to relate a response variable $Y$ to a set of potential explanatory variables $X = (X_1,\dots,X_p)$, and start by trying to identify variables that contribute to this relationship. In statistical…

Statistics Theory · Mathematics 2020-10-07 Wenshuo Wang , Lucas Janson

Variable selection for structured covariates lying on an underlying known graph is a problem motivated by practical applications, and has been a topic of increasing interest. However, most of the existing methods may not be scalable to high…

Methodology · Statistics 2016-04-27 Changgee Chang , Suprateek Kundu , Qi Long

The recent paper Cand\`es et al. (2018) introduced model-X knockoffs, a method for variable selection that provably and non-asymptotically controls the false discovery rate with no restrictions or assumptions on the dimensionality of the…

Methodology · Statistics 2020-06-16 Dongming Huang , Lucas Janson

The scan statistic is widely used in spatial cluster detection applications of inhomogeneous Poisson processes. However, real data may present substantial departure from the underlying Poisson process. One of the possible departures has to…

Methodology · Statistics 2013-11-19 André L. F. Cançado , Cibele Q. da-Silva , Michel F. da Silva

Controlling the False Discovery Rate (FDR) is critical for reproducible variable selection, especially given the prevalence of complex predictive modeling. The recent Split Knockoff method, an extension of the canonical Knockoffs framework,…

Methodology · Statistics 2025-09-05 Yang Cao , Hangyu Lin , Xinwei Sun , Yuan Yao

The knockoff-based multiple testing setup of Barber & Candes (2015) for variable selection in multiple regression where sample size is as large as the number of explanatory variables is considered. The method of Benjamini & Hochberg (1995)…

Methodology · Statistics 2021-08-20 Sanat K. Sarkar , Cheng Yong Tang

Single-cell sequencing technologies have significantly advanced molecular and cellular biology, offering unprecedented insights into cellular heterogeneity by allowing for the measurement of gene expression at an individual cell level.…

Methodology · Statistics 2024-03-26 Junsouk Choi , Hee Cheol Chung , Irina Gaynanova , Yang Ni
‹ Prev 1 3 4 5 6 7 10 Next ›