English
Related papers

Related papers: How does limma-trend work? An empirical partially …

200 papers

In recent years, multiple hypothesis testing has come to the forefront of statistical research, ostensibly in relation to applications in genomics and some other emerging fields. The false discovery rate (FDR) and its variants provide very…

Statistics Theory · Mathematics 2008-12-18 Subhashis Ghosal , Anindya Roy , Yongqiang Tang

When testing a number of statistical hypotheses using data from location families, it is often useful to control the false discovery rate (FDR) not just for hypotheses of the null values but also of other parameter values that are deemed…

Methodology · Statistics 2026-05-12 Zijun Gao , Wenjie Hu , Qingyuan Zhao

As LLMs continue to scale, improving training efficiency increasingly depends on using data more effectively. Data selection addresses this problem by allocating a limited training budget to samples that best promote a target behavior.…

Machine Learning · Computer Science 2026-05-21 Qihao Lin , Guanxu Chen , Dongrui Liu , Jing Shao

In high dimensional variable selection problems, statisticians often seek to design multiple testing procedures that control the False Discovery Rate (FDR), while concurrently identifying a greater number of relevant variables. Model-X…

Statistics Theory · Mathematics 2023-07-25 Taejoo Ahn , Licong Lin , Song Mei

In the high dimensional regression analysis when the number of predictors is much larger than the sample size, an important question is to select the important variable which are relevant to the response variable of interest. Variable…

Methodology · Statistics 2023-01-09 Pengsheng Ji , Zhigen Zhao

Bayesian inference provides a flexible way of combining data with prior information. However, quantile regression is not equipped with a parametric likelihood, and therefore, Bayesian inference for quantile regression demands careful…

Statistics Theory · Mathematics 2012-07-24 Yunwen Yang , Xuming He

Controlling the false discovery rate (FDR) is a critical challenge in large-scale data analysis, particularly in the presence of outliers. A common practice involves imposing a Student-$t$ distribution to eliminate the influence of…

Methodology · Statistics 2026-04-01 Yoshiko Hayashi

Random forests is a state-of-the-art supervised machine learning method which behaves well in high-dimensional settings although some limitations may happen when $p$, the number of predictors, is much larger than the number of observations…

Methodology · Statistics 2019-02-01 Louis Capitaine , Robin Genuer , Rodolphe Thiébaut

Given a random sample from a distribution with density function that depends on an unknown parameter $\theta$, we are interested in accurately estimating the true parametric density function at a future observation from the same…

Statistics Theory · Mathematics 2009-09-29 Mihaela Aslan

Factor models are widely used for dimension reduction. Bayesian approaches to these models often place a prior on the factor loadings that allows for infinitely many factors, with loadings increasingly shrunk toward zero as the column index…

Methodology · Statistics 2026-03-31 Shicheng Liu , Qingping Zhou , Yanan Fan , Xiongwen Ke

The models used to describe the kinetics of ruminal degradation are usually nonlinear models where the dependent variable is the proportion of degraded food. The method of least squares is the standard approach used to estimate the unknown…

Applications · Statistics 2021-07-21 Diego Salmerón

We analyze the posterior contraction rates of parameters in Bayesian models via the Langevin diffusion process, in particular by controlling moments of the stochastic process and taking limits. Analogous to the non-asymptotic analysis of…

Statistics Theory · Mathematics 2022-08-18 Wenlong Mou , Nhat Ho , Martin J. Wainwright , Peter Bartlett , Michael I. Jordan

Modern data analysis frequently involves large-scale hypothesis testing, which naturally gives rise to the problem of maintaining control of a suitable type I error rate, such as the false discovery rate (FDR). In many biomedical and…

Methodology · Statistics 2023-07-25 David S. Robertson , James M. S. Wason , Aaditya Ramdas

Multiple testing literature contains ample research on controlling false discoveries for hypotheses classified according to one criterion, which we refer to as one-way classified hypotheses. Although simultaneous classification of…

Methodology · Statistics 2019-03-12 Shinjini Nandi , Sanat K. Sarkar

We present a new approach to semiparametric inference using corrected posterior distributions. The method allows us to leverage the adaptivity, regularization and predictive power of nonparametric Bayesian procedures to estimate…

Methodology · Statistics 2023-06-21 Andrew Yiu , Edwin Fong , Chris Holmes , Judith Rousseau

We provide new non-asymptotic false discovery proportion (FDP) confidence envelopes in several multiple testing settings relevant for modern high dimensional-data methods. We revisit the multiple testing scenarios considered in the recent…

Statistics Theory · Mathematics 2024-09-18 Iqraa Meah , Gilles Blanchard , Etienne Roquain

Existing pretraining data mixing methods for large language models (LLMs) typically follow a domain-wise methodology, a top-down process that first determines domain weights and then performs uniform data sampling across each domain.…

Computation and Language · Computer Science 2025-03-04 Xiangyu Xi , Deyang Kong , Jian Yang , Jiawei Yang , Zhengyu Chen , Wei Wang , Jingang Wang , Xunliang Cai , Shikun Zhang , Wei Ye

Modern scientific technology has provided a new class of large-scale simultaneous inference problems, with thousands of hypothesis tests to consider at the same time. Microarrays epitomize this type of technology, but similar situations…

Statistics Theory · Mathematics 2007-11-06 Bradley Efron

Epidemiologic screening programs often make use of tests with small, but non-zero probabilities of misdiagnosis. In this article, we assume the target population is finite with a fixed number of true cases, and that we apply an imperfect…

Methodology · Statistics 2024-04-22 Lin Ge , Yuzi Zhang , Lance A. Waller , Robert H. Lyles

Features in predictive models are not exchangeable, yet common supervised models treat them as such. Here we study ridge regression when the analyst can partition the features into $K$ groups based on external side-information. For example,…

Methodology · Statistics 2021-03-05 Nikolaos Ignatiadis , Panagiotis Lolas