English
Related papers

Related papers: Sample Size Determination Under Selection Bias: Ro…

200 papers

A common task in high-throughput biology is to test for differences in means between two samples across thousands of features (e.g., genes or proteins), often with only a handful of replicates per sample. Moderated t-tests handle this…

Methodology · Statistics 2025-10-02 Wanyi Ling , Wufang Hong , Nikolaos Ignatiadis

The finite sensitivity of instruments or detection methods means that data sets in many areas of astronomy, for example cosmological or exoplanet surveys, are necessarily systematically incomplete. Such data sets, where the population being…

Instrumentation and Methods for Astrophysics · Physics 2020-10-14 Adam B. Mantz

Historically, to bound the mean for small sample sizes, practitioners have had to choose between using methods with unrealistic assumptions about the unknown distribution (e.g., Gaussianity) and methods like Hoeffding's inequality that use…

Statistics Theory · Mathematics 2021-10-27 My Phan , Philip S. Thomas , Erik Learned-Miller

Sparseness of the regression coefficient vector is often a desirable property, since, among other benefits, sparseness improves interpretability. In practice, many true regression coefficients might be negligibly small, but non-zero, which…

Methodology · Statistics 2019-10-01 Daniel Andrade , Kenji Fukumizu

Infinite-order U-statistics (IOUS) has been used extensively on subbagging ensemble learning algorithms such as random forests to quantify its uncertainty. While normality results of IOUS have been studied extensively, its variance…

Machine Learning · Statistics 2023-02-16 Tianning Xu , Ruoqing Zhu , Xiaofeng Shao

Data-driven risk analysis involves the inference of probability distributions from measured or simulated data. In the case of a highly reliable system, such as the electricity grid, the amount of relevant data is often exceedingly limited,…

Methodology · Statistics 2017-07-11 Simon H. Tindemans , Goran Strbac

In this paper, we address the challenge of sampling in scenarios where limited resources prevent exhaustive measurement across all subjects. We consider a setting where samples are drawn from multiple groups, each following a distribution…

Econometrics · Economics 2024-08-29 Carol Liu

Probability proportional to size (PPS) sampling schemes with a target sample size aim to produce a sample comprising a specified number $n$ of items while ensuring that each item in the population appears in the sample with a probability…

Methodology · Statistics 2024-11-14 Brian Hentschel , Peter J. Haas , Yuanyuan Tian

We describe the distribution of frequencies ordered by sample values in a random sample of size $n$ from the two parameter GEM$(\alpha,\theta)$ random discrete distribution on the positive integers. These frequencies are a…

Probability · Mathematics 2017-08-29 Jim Pitman , Yuri Yakubovich

In this paper, we establish the central limit theorem (CLT) for linear spectral statistics (LSS) of large-dimensional sample covariance matrix when the population covariance matrices are not uniformly bounded, which is a nontrivial…

Statistics Theory · Mathematics 2022-05-17 Zhijun Liu , Jiang Hu , Zhidong Bai , Haiyan Song

To learn (statistical) dependencies among random variables requires exponentially large sample size in the number of observed random variables if any arbitrary joint probability distribution can occur. We consider the case that sparse data…

Machine Learning · Computer Science 2007-05-23 Dominik Janzing , Daniel Herrmann

Biomedical researchers usually study the effects of certain exposures on disease risks among a well-defined population. To achieve this goal, the gold standard is to design a trial with an appropriate sample from that population. Due to the…

Applications · Statistics 2019-11-18 Cheng Zheng , Sayan Dasgupta , Yuxiang Xie , Asad Haris , Ying Qing Chen

In a completely randomized experiment, the variances of treatment effect estimators in the finite population are usually not identifiable and hence not estimable. Although some estimable bounds of the variances have been established in the…

Statistics Theory · Mathematics 2022-09-20 Ruoyu Wang , Qihua Wang , Wang Miao , Xiaohua Zhou

Statistical learning theory chiefly studies restricted hypothesis classes, particularly those with finite Vapnik-Chervonenkis (VC) dimension. The fundamental quantity of interest is the sample complexity: the number of samples required to…

Machine Learning · Computer Science 2008-07-10 David Soloveichik

Nonprobability (convenience) samples are increasingly sought to stabilize estimations for one or more population variables of interest that are performed using a randomized survey (reference) sample by increasing the effective sample size.…

A new framework is introduced for examining and evaluating the fundamental limits of lossless data compression, that emphasizes genuinely non-asymptotic results. The {\em sample complexity} of compressing a given source is defined as the…

Information Theory · Computer Science 2026-04-16 Terence Viaud , Ioannis Kontoyiannis

Astronomers are often confronted with funky populations and distributions of objects: brighter objects are more likely to be detected; targets are selected based on colour cuts; imperfect classification yields impure samples. Failing to…

Cosmology and Nongalactic Astrophysics · Physics 2017-06-21 Samuel R. Hinton , Alex Kim , Tamara M. Davis

Self-training is a well-known approach for semi-supervised learning. It consists of iteratively assigning pseudo-labels to unlabeled data for which the model is confident and treating them as labeled examples. For neural networks, softmax…

Machine Learning · Computer Science 2024-04-04 Ambroise Odonnat , Vasilii Feofanov , Ievgen Redko

A cornerstone of the classical view of tolerance is the elimination of self-reactive T cells during negative selection in the thymus. However, high-throughput T-cell receptor sequencing data has so far failed to detect substantial…

Populations and Evolution · Quantitative Biology 2023-03-14 Thierry Mora , Aleksandra M. Walczak

We propose a cautious Bayesian variable selection routine by investigating the sensitivity of a hierarchical model, where the regression coefficients are specified by spike and slab priors. We exploit the use of latent variables to…

Methodology · Statistics 2022-06-20 Tathagata Basu , Matthias C. M. Troffaes , Jochen Einbeck