English
Related papers

Related papers: Effect of pooling samples on the efficiency of com…

200 papers

Many large-scale testing procedures learn signal structure from the data to boost power. Direct data reuse can inflate Type-I error ("double dipping"), so a common remedy is masking: withholding some information during learning and using it…

Statistics Theory · Mathematics 2026-04-02 Abhinav Chakraborty , Junu Lee , Eugene Katsevich

In randomized experiments with noncompliance, tests may focus on compliers rather than on the overall sample. Rubin (1998) put forth such a method, and argued that testing for the complier average causal effect and averaging permutation…

Methodology · Statistics 2016-02-23 Laura Forastiere , Fabrizia Mealli , Luke Miratrix

Modern scientific technology has provided a new class of large-scale simultaneous inference problems, with thousands of hypothesis tests to consider at the same time. Microarrays epitomize this type of technology, but similar situations…

Statistics Theory · Mathematics 2007-11-06 Bradley Efron

Calculating free energy differences is a topic of substantial interest and has many applications including molecular docking and hydration, solvation, and binding free energies which is used in computational drug discovery. However, in…

Chemical Physics · Physics 2013-10-16 Asaf Farhi

Extrapolating treatment effects from related studies is a promising strategy for designing and analyzing clinical trials in situations where achieving an adequate sample size is challenging. Bayesian methods are well-suited for this…

Methodology · Statistics 2025-11-25 Tristan Fauvel , Julien Tanniou , Pascal Godbillot , Marie Génin , Billy Amzal

Free electrons in the interstellar medium cause frequency-dependent delays in pulse arrival times due to both scattering and dispersion. Multi-frequency measurements are used to estimate and remove dispersion delays. In this paper, we focus…

High Energy Astrophysical Phenomena · Physics 2015-06-23 M. T. Lam , J. M. Cordes , S. Chatterjee , T. Dolch

This paper shows how to use a randomized saturation experimental design to identify and estimate causal effects in the presence of spillovers--one person's treatment may affect another's outcome--and one-sided non-compliance--subjects can…

Machine Learning methods have of late made significant efforts to solving multidisciplinary problems in the field of cancer classification using microarray gene expression data. Feature subset selection methods can play an important role in…

Computational Engineering, Finance, and Science · Computer Science 2013-03-04 G. Prat , Ll. Belanche

Phylogenetic comparative methods may fail to produce meaningful results when either the underlying model is inappropriate or the data contain insufficient information to inform the inference. The ability to measure the statistical power of…

Quantitative Methods · Quantitative Biology 2012-07-26 Carl Boettiger , Graham Coop , Peter Ralph

Participant level meta-analysis across multiple studies increases the sample size for pooled analyses, thereby improving precision in effect estimates and enabling subgroup analyses. For analyses involving biomarker measurements as an…

Bibliometricians face several issues when drawing and analyzing samples of citation records for their research. Drawing samples that are too small may make it difficult or impossible for studies to achieve their goals, while drawing samples…

Applications · Statistics 2015-11-18 Richard Williams , Lutz Bornmann

Detection of rare traits or diseases in a large population is challenging. Pool testing allows covering larger swathes of population at a reduced cost, while simplifying logistics. However, testing precision decreases as it becomes unclear…

Information Theory · Computer Science 2021-06-22 Éric Brier , Megi Dervishi , Rémi Géraud-Stewart , David Naccache , Ofer Yifrach-Stav

As deep learning based models are increasingly being used for information retrieval (IR), a major challenge is to ensure the availability of test collections for measuring their quality. Test collections are generated based on pooling…

Information Retrieval · Computer Science 2020-04-29 Emine Yilmaz , Nick Craswell , Bhaskar Mitra , Daniel Campos

Basket trials are increasingly used for the simultaneous evaluation of a new treatment in various patient subgroups under one overarching protocol. We propose a Bayesian approach to sample size determination in basket trials that permit…

Methodology · Statistics 2022-09-02 Haiyan Zheng , Michael J. Grayling , Pavel Mozgunov , Thomas Jaki , James M. S. Wason

Linear combinations of multinomial probabilities, such as those resulting from contingency tables, are of use when evaluating classification system performance. While large sample inference methods for these combinations exist, small sample…

Methodology · Statistics 2021-04-20 Katherine A. Batterton , Christine M. Schubert , Richard L. Warr

Composite binary endpoints are increasingly used as primary endpoints in clinical trials. When designing a trial, it is crucial to determine the appropriate sample size for testing the statistical differences between treatment groups for…

Applications · Statistics 2019-01-15 Marta Bofill Roig , Guadalupe Gómez Melis

Pooled testing is a common strategy for public health disease screening under limited testing resources, allowing multiple biological samples to be tested together with the resources of a single test, at the cost of reduced individual…

Computer Science and Game Theory · Computer Science 2026-02-02 Nicholas Lopez , Francisco Marmolejo-Cossío , Jose Roberto Tello Ayala , David C. Parkes

Large-scale testing is crucial in pandemic containment, but resources are often prohibitively constrained. We study the optimal application of pooled testing for populations that are heterogeneous with respect to an individual's infection…

Computer Science and Game Theory · Computer Science 2023-09-22 Simon Finster , Michelle González Amador , Edwin Lock , Francisco Marmolejo-Cossío , Evi Micha , Ariel D. Procaccia

Causal discovery can be a powerful tool for investigating causality when a system can be observed but is inaccessible to experiments in practice. Despite this, it is rarely used in any scientific or medical fields. One of the major hurdles…

Machine Learning · Statistics 2019-10-07 Erich Kummerfeld , Alexander Rix

The choice of sample size in the context of co-primary endpoints for a randomised trial is discussed. Current guidance can leave endpoints with unequal marginal power. A method is provided to achieve equal marginal power by using the…

Methodology · Statistics 2026-02-23 Simon Bond