English
Related papers

Related papers: ECKO: Ensemble of Clustered Knockoffs for multivar…

200 papers

With the development of data collection techniques, analysis with a survival response and high-dimensional covariates has become routine. Here we consider an interaction model, which includes a set of low-dimensional covariates, a set of…

Methodology · Statistics 2023-11-27 Weijuan Liang , Qingzhao Zhang , Shuangge Ma

Clustering is a central tool in biomedical research for discovering heterogeneous patient subpopulations, where group boundaries are often diffuse rather than sharply separated. Traditional methods produce hard partitions, whereas soft…

Methodology · Statistics 2026-01-07 Qiuyi Wu , Zihan Zhu , Anru R. Zhang

Motivation: In microarray analysis, special consideration must be given to the issues of multiple statistical tests and typically p-values are adjusted to control family-wise error rate (FWER) or false discovery rate (FDR). FDR metrics have…

Quantitative Methods · Quantitative Biology 2007-05-23 Rishi L. Khan , Rajanikanth Vadigepalli , Guang Gao , James S. Schwaber

In large scale multiple testing problems, a two-class empirical Bayes approach can be used to control the false discovery rate (Fdr) for the entire array of hypotheses under study. A sample splitting step is incorporated to modify that…

Computation · Statistics 2019-12-13 Paramita Chakraborty , Chong Ma , John Grego , James Lynch

Out of the participants in a randomized experiment with anticipated heterogeneous treatment effects, is it possible to identify which subjects have a positive treatment effect? While subgroup analysis has received attention, claims about…

Methodology · Statistics 2024-05-14 Boyan Duan , Larry Wasserman , Aaditya Ramdas

We present false discovery rate smoothing, an empirical-Bayes method for exploiting spatial structure in large multiple-testing problems. FDR smoothing automatically finds spatially localized regions of significant test statistics. It then…

Methodology · Statistics 2016-11-15 Wesley Tansey , Oluwasanmi Koyejo , Russell A. Poldrack , James G. Scott

Competition-based approach to controlling the false discovery rate (FDR) recently rose to prominence when, generalizing it to sequential hypothesis testing, Barber and Cand\`es used it as part of their knockoff-filter. Control of the FDR…

Methodology · Statistics 2023-02-24 Arya Ebadi , Dong Luo , Jack Freestone , William Stafford Noble , Uri Keich

Multimodal learning aims to enhance perceptual and decision-making capabilities by integrating information from diverse sources. However, classical deep learning approaches face a critical trade-off between the high accuracy of black-box…

Quantum Physics · Physics 2026-01-14 Yu Wu , Qianli Zhou , Jie Geng , Xinyang Deng , Wen Jiang

Analyzing large-scale, multi-experiment studies requires scientists to test each experimental outcome for statistical significance and then assess the results as a whole. We present Black Box FDR (BB-FDR), an empirical-Bayes method for…

Machine Learning · Statistics 2018-06-11 Wesley Tansey , Yixin Wang , David M. Blei , Raul Rabadan

We propose a novel, succinct, and effective approach for distribution prediction to quantify uncertainty in machine learning. It incorporates adaptively flexible distribution prediction of $\mathbb{P}(\mathbf{y}|\mathbf{X}=x)$ in regression…

Machine Learning · Computer Science 2023-06-21 Xing Yan , Yonghua Su , Wenxuan Ma

We study the high-dimensional linear regression problem with categorical predictors that have many levels. We propose a new estimation approach, which performs model compression via two mechanisms by simultaneously encouraging (a)…

Methodology · Statistics 2026-03-30 Kayhan Behdin , Riade Benbaki , Peter Radchenko , Rahul Mazumder

Reliable detection and quantification of quantum entanglement, particularly in high-spin or many-body systems, present significant computational challenges for traditional methods. This study examines the effectiveness of ensemble machine…

Quantum Physics · Physics 2025-07-18 M. Y. Abd-Rabbou , Amr M. Abdallah , Ahmed A. Zahia , Ashraf A. Gouda , Cong-Feng Qiao

A growth curve model (GCM) aims to characterize how an outcome variable evolves, develops and grows as a function of time, along with other predictors. It provides a particularly useful framework to model growth trend in longitudinal data.…

Methodology · Statistics 2023-12-29 Xin Zhou , Yin Xia , Lexin Li

In real-world applications of reinforcement learning, it is often challenging to obtain a state representation that is parsimonious and satisfies the Markov property without prior knowledge. Consequently, it is common practice to construct…

Machine Learning · Statistics 2024-07-31 Tao Ma , Jin Zhu , Hengrui Cai , Zhengling Qi , Yunxiao Chen , Chengchun Shi , Eric B. Laber

Many community detection algorithms are inherently stochastic, leading to variations in their output depending on input parameters and random seeds. This variability makes the results of a single run of these algorithms less reliable.…

Social and Information Networks · Computer Science 2025-02-25 Yasamin Tabatabaee , Eleanor Wedell , Minhyuk Park , Tandy Warnow

Clustering is a fundamental unsupervised learning approach. Many clustering algorithms -- such as $k$-means -- rely on the euclidean distance as a similarity measure, which is often not the most relevant metric for high dimensional data…

Machine Learning · Computer Science 2019-10-22 Aude Genevay , Gabriel Dulac-Arnold , Jean-Philippe Vert

We apply the knockoff procedure to factor selection in finance. By building fake but realistic factors, this procedure makes it possible to control the fraction of false discovery in a given set of factors. To show its versatility, we apply…

Statistical Finance · Quantitative Finance 2021-07-07 Damien Challet , Christian Bongiorno , Guillaume Pelletier

The generalized linear models (GLM) have been widely used in practice to model non-Gaussian response variables. When the number of explanatory features is relatively large, scientific researchers are of interest to perform controlled…

Methodology · Statistics 2020-07-03 Chenguang Dai , Buyu Lin , Xin Xing , Jun S. Liu

Cluster inference based on spatial extent thresholding is the most popular analysis method for finding activated brain areas in neuroimaging. However, the method has several well-known issues. While powerful for finding brain regions with…

Methodology · Statistics 2022-08-10 Jelle J. Goeman , Paweł\ Górecki , Ramin Monajemi , Xu Chen , Thomas E. Nichols , Wouter Weeda

High-dimensional clustering analysis is a challenging problem in statistics and machine learning, with broad applications such as the analysis of microarray data and RNA-seq data. In this paper, we propose a new clustering procedure called…

Methodology · Statistics 2022-10-31 Tianqi Liu , Yu Lu , Biqing Zhu , Hongyu Zhao
‹ Prev 1 8 9 10 Next ›