English
Related papers

Related papers: Beyond the E-value: stratified statistics for prot…

200 papers

Revealing the functional sites of biological sequences, such as evolutionary conserved, structurally interacting or co-evolving protein sites, is a fundamental, and yet challenging task. Different frameworks and models were developed to…

Quantitative Methods · Quantitative Biology 2019-06-07 Justas Dauparas , Haobo Wang , Avi Swartz , Peter Koo , Mor Nitzan , Sergey Ovchinnikov

Delayed separation of survival curves is a common occurrence in confirmatory studies in immuno-oncology. Many novel statistical methods that aim to efficiently capture potential long-term survival improvements have been proposed in recent…

Methodology · Statistics 2022-01-26 Dominic Magirr , José L. Jiménez

We study the sample efficiency of domain randomization and robust control for the benchmark problem of learning the linear quadratic regulator (LQR). Domain randomization, which synthesizes controllers by minimizing average performance over…

Systems and Control · Electrical Eng. & Systems 2025-02-19 Tesshu Fujinami , Bruce D. Lee , Nikolai Matni , George J. Pappas

Multiple hypotheses testing is a core problem in statistical inference and arises in almost every scientific field. Given a sequence of null hypotheses $\mathcal{H}(n) = (H_1,..., H_n)$, Benjamini and Hochberg…

Methodology · Statistics 2015-03-05 Adel Javanmard , Andrea Montanari

Stratification in both the design and analysis of randomized clinical trials is common. Despite features in automated randomization systems to re-confirm the stratifying variables, incorrect values of these variables may be entered. These…

Methodology · Statistics 2023-07-24 Neal Thomas

Randomization, as a key technique in clinical trials, can eliminate sources of bias and produce comparable treatment groups. In randomized experiments, the treatment effect is a parameter of general interest. Researchers have explored the…

Methodology · Statistics 2023-12-05 Fuyi Tu , Wei Ma , Hanzhong Liu

Protein function prediction is currently achieved by encoding its sequence or structure, where the sequence-to-function transcendence and high-quality structural data scarcity lead to obvious performance bottlenecks. Protein domains are…

Biomolecules · Quantitative Biology 2024-12-03 Mingqing Wang , Zhiwei Nie , Yonghong He , Athanasios V. Vasilakos , Zhixiang Ren

Analyzing time series in the frequency domain enables the development of powerful tools for investigating the second-order characteristics of multivariate processes. Parameters like the spectral density matrix and its inverse, the coherence…

Methodology · Statistics 2024-01-19 Jonas Krampe , Efstathios Paparoditis

Domain adaptation aims to mitigate performance degradation caused by distribution shifts between a labeled source domain and an unlabeled or sparsely labeled target domain. Most existing approaches estimate domain discrepancy either in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Xi Ding , Lei Wang , Syuan-Hao Li , Yongsheng Gao

In biomedical studies it is of substantial interest to develop risk prediction scores using high-dimensional data such as gene expression data for clinical endpoints that are subject to censoring. In the presence of well-established…

Applications · Statistics 2011-11-24 Qi Long , Matthias Chung , Carlos S. Moreno , Brent A. Johnson

Variable selection has been widely used in data analysis for the past decades, and it becomes increasingly important in the Big Data era as there are usually hundreds of variables available in a dataset. To enhance interpretability of a…

Methodology · Statistics 2020-08-17 Yuxiang Xie , Kwun Chuen Gary Chan

Distribution Regression (DR) on stochastic processes describes the learning task of regression on collections of time series. Path signatures, a technique prevalent in stochastic analysis, have been used to solve the DR problem. Recent…

Machine Learning · Computer Science 2024-10-15 Andrew Alden , Carmine Ventre , Blanka Horvath

The False Discovery Rate (FDR) method has recently been described by Miller et al (2001), along with several examples of astrophysical applications. FDR is a new statistical procedure due to Benjamini and Hochberg (1995) for controlling the…

Astrophysics · Physics 2009-11-07 A. M. Hopkins , C. J. Miller , A. J. Connolly , C. Genovese , R. C. Nichol , L. Wasserman

Empirical substitution matrices represent the average tendencies of substitutions over various protein families by sacrificing gene-level resolution. We develop a codon-based model, in which mutational tendencies of codon, a genetic code,…

Populations and Evolution · Quantitative Biology 2011-08-31 Sanzo Miyazawa

Improved procedures, in terms of smaller missed discovery rates (MDR), for performing multiple hypotheses testing with weak and strong control of the family-wise error rate (FWER) or the false discovery rate (FDR) are developed and studied.…

Statistics Theory · Mathematics 2011-03-10 Edsel A. Peña , Joshua D. Habiger , Wensong Wu

In many large scale multiple testing applications, the hypotheses often have a known graphical structure, such as gene ontology in gene expression data. Exploiting this graphical structure in multiple testing procedures can improve power as…

Methodology · Statistics 2018-12-04 Wenge Guo , Gavin Lynch , Joseph P. Romano

Simultaneously performing variable selection and inference in high-dimensional regression models is an open challenge in statistics and machine learning. The increasing availability of vast amounts of variables requires the adoption of…

Methodology · Statistics 2025-05-08 Marco Molinari , Magne Thoresen

Post-treatment variables often complicate causal inference. They appear in many scientific problems, including noncompliance, truncation by death, mediation, and surrogate endpoint evaluation. Principal stratification is a strategy to…

Methodology · Statistics 2024-04-04 Sizhu Lu , Zhichao Jiang , Peng Ding

We address the multiple testing problem under the assumption that the true/false hypotheses are driven by a Hidden Markov Model (HMM), which is recognized as a fundamental setting to model multiple testing under dependence since the seminal…

Methodology · Statistics 2021-05-04 Marie Perrot-Dockès , Gilles Blanchard , Pierre Neuvial , Etienne Roquain

The problem of validating or criticising models for georeferenced data is challenging, since the conclusions can vary significantly depending on the locations of the validation set. This work proposes the use of cross-validation techniques…

Computation · Statistics 2018-02-19 Viviana G R Lobo , Thaís C O da Fonseca , Fernando A S Moura