English
Related papers

Related papers: Beyond the E-value: stratified statistics for prot…

200 papers

Inequalities are key tools to prove FDR control of a multiple test. The present paper studies upper and lower bounds for the FDR under various dependence structures of p-values, namely independence, reverse martingale dependence and…

Statistics Theory · Mathematics 2015-02-18 Philipp Heesen , Arnold Janssen

Sufficient dimension reduction (SDR) methods, which often rely on class precision matrices, are widely used in supervised statistical classification problems. However, when class-specific sample sizes are small relative to the original…

Methodology · Statistics 2025-06-25 Derik T. Boonstra , Rakheon Kim , Dean M. Young

Controlling false discovery rate (FDR) while leveraging the side information of multiple hypothesis testing is an emerging research topic in modern data science. Existing methods rely on the test-level covariates while ignoring possible…

Machine Learning · Statistics 2021-01-26 Lin Qiu , Nils Murrugarra-Llerena , Vítor Silva , Lin Lin , Vernon M. Chinchilli

Multiple hypothesis testing, a situation when we wish to consider many hypotheses, is a core problem in statistical inference that arises in almost every scientific field. In this setting, controlling the false discovery rate (FDR), which…

Statistics Theory · Mathematics 2019-03-19 Shiyun Chen , Shiva Kasiviswanathan

Recent computational advances in the accurate prediction of protein three-dimensional (3D) structures from amino acid sequences now present a unique opportunity to decipher the interrelationships between proteins. This task entails--but is…

Biomolecules · Quantitative Biology 2020-05-19 Menuka Jaiswal , Saad Saleem , Yonghyeon Kweon , Eli J Draizen , Stella Veretnik , Cameron Mura , Philip E. Bourne

Protein structure prediction is one of the most important problems in computational biology. The most successful computational approach, also called template-based modeling, identifies templates with solved crystal structures for the query…

Biomolecules · Quantitative Biology 2013-06-20 Jian Peng

There has been recent interest in extending the ideas of False Discovery Rates (FDR) to variable selection in regression settings. Traditionally the FDR in these settings has been defined in terms of the coefficients of the full regression…

Methodology · Statistics 2013-02-12 Max Grazier G'Sell , Trevor Hastie , Robert Tibshirani

Estimating local false discovery rates (fdr) is central to large-scale multiple hypothesis testing, yet different methods often produce divergent results, and there is little guidance for selecting among them. Because ground truth…

Methodology · Statistics 2025-09-17 Jenna M. Landy , Giovanni Parmigiani

For time series arising from latent dynamical systems, existing cross-domain generalization methods commonly assume that samples are comparably meaningful within a shared representation space. In real-world settings, however, different…

Machine Learning · Computer Science 2026-03-04 Jinyang Li , Shuhao Mei , Xiaoyu Xiao , Shuhang Li , Ruoxi Yun , Jinbo Sun

Semiparametric statistics play a pivotal role in a wide range of domains, including but not limited to missing data, causal inference, and transfer learning, to name a few. In many settings, semiparametric theory leads to (nearly)…

Machine Learning · Statistics 2024-08-06 Qinshuo Liu , Zixin Wang , Xi-An Li , Xinyao Ji , Lei Zhang , Lin Liu , Zhonghua Liu

The accurate prediction of changes in protein stability under multiple amino acid substitutions is essential for realising true in-silico protein re-design. To this purpose, we propose improvements to state-of-the-art Deep learning (DL)…

Biomolecules · Quantitative Biology 2023-06-01 Sebastien Boyer , Sam Money-Kyrle , Oliver Bent

By restricting the possible values of the proportion of null hypotheses that are true, the local false discovery rate (LFDR) can be estimated using as few as one comparison. The proportion of proteins with equivalent abundance was estimated…

Methodology · Statistics 2011-05-13 David R. Bickel

Sequential decision making significantly speeds up research and is more cost-effective compared to fixed-n methods. We present a method for sequential decision making for stratified count data that retains Type-I error guarantee or false…

Methodology · Statistics 2023-02-23 Rosanne J. Turner , Peter D. Grünwald

Quantifying how genomic features influence different parts of an outcome distribution requires statistical tools that go beyond mean regression, especially in ultrahigh-dimensional settings. Motivated by the study of LINE-1 activity in…

Methodology · Statistics 2025-11-27 Sang Kyu Lee , Tongwu Zhang , Hyokyoung G. Hong , Haolei Weng

This paper develops a unified framework for partial identification and inference in stratified experiments with attrition, accommodating both equal and heterogeneous treatment shares across strata. For equal-share designs, we apply recent…

Econometrics · Economics 2026-01-21 Bruno Ferman , Davi Siqueira , Vitor Possebom

We present a novel necessary and sufficient principle for False Discovery Rate (FDR) control. This e-Partitioning Principle says that a procedure controls FDR if and only if it is a special case of a general e-Partitioning procedure. By…

Statistics Theory · Mathematics 2025-09-15 Jelle Goeman , Rianne de Heide , Aldo Solari

This paper outlines a unified framework for high dimensional variable selection for classification problems. Traditional approaches to finding interesting variables mostly utilize only partial information through moments (like mean…

Methodology · Statistics 2016-11-25 S. Mukhopadhyay , Emanuel Parzen , S. N. Lahiri

Large-scale multiple testing with highly correlated test statistics arises frequently in many scientific research. Incorporating correlation information in estimating false discovery proportion has attracted increasing attention in recent…

Methodology · Statistics 2019-03-28 Jianqing Fan , Xu Han

Controlling the false discovery rate (FDR) in variable selection becomes challenging when predictors are correlated, as existing methods often exclude all members of correlated groups and consequently perform poorly for prediction. We…

Methodology · Statistics 2026-03-03 Sarah Organ , Toby Kenney , Hong Gu

E-values have gained attention as potential alternatives to p-values as measures of uncertainty, significance and evidence. In brief, e-values are realized by random variables with expectation at most one under the null; examples include…

Statistics Theory · Mathematics 2021-12-16 Ruodu Wang , Aaditya Ramdas