English
Related papers

Related papers: Designed Sampling from Large Databases for Control…

200 papers

Pre-experiment stratification, or blocking, is a well-established technique for designing more efficient experiments and increasing the precision of the experimental estimates. However, when researchers have access to many covariates at the…

Econometrics · Economics 2025-10-01 George Gui , Seungwoo Kim

An important task in drug development is to identify patients, which respond better or worse to an experimental treatment. Identifying predictive covariates, which influence the treatment effect and can be used to define subgroups of…

Methodology · Statistics 2018-11-27 Marius Thomas , Björn Bornkamp , Katja Ickstadt

Questions of `how best to acquire data' are essential to modeling and prediction in the natural and social sciences, engineering applications, and beyond. Optimal experimental design (OED) formalizes these questions and creates…

Methodology · Statistics 2026-05-01 Xun Huan , Jayanth Jagalur , Youssef Marzouk

The design of multiple experiments is commonly undertaken via suboptimal strategies, such as batch (open-loop) design that omits feedback or greedy (myopic) design that does not account for future effects. This paper introduces new…

Methodology · Statistics 2016-04-29 Xun Huan , Youssef M. Marzouk

We propose Distributionally Balanced Designs (DBD), a new class of probability sampling designs that target representativeness at the level of the full auxiliary distribution rather than selected moments. In disciplines such as ecology,…

Methodology · Statistics 2026-03-13 Anton Grafström , Wilmer Prentius

Estimating how a treatment affects units individually, known as heterogeneous treatment effect (HTE) estimation, is an essential part of decision-making and policy implementation. The accumulation of large amounts of data in many domains,…

Machine Learning · Computer Science 2022-06-28 Christopher Tran , Elena Zheleva

Observational studies often benefit from an abundance of observational units. This can lead to studies that -- while challenged by issues of internal validity -- have inferences derived from sample sizes substantially larger than randomized…

Methodology · Statistics 2020-08-24 Rachael C. Aikens , Dylan Greaves , Michael Baiocchi

Randomized controlled trials (RCTs) can be used to generate guarantees on treatment effects. However, RCTs often spend unnecessary resources exploring sub-optimal treatments, which can reduce the power of treatment guarantees. To address…

Computers and Society · Computer Science 2024-10-16 Santiago Cortes-Gomez , Naveen Raman , Aarti Singh , Bryan Wilder

Indirect experiments provide a valuable framework for estimating treatment effects in situations where conducting randomized control trials (RCTs) is impractical or unethical. Unlike RCTs, indirect experiments estimate treatment effects by…

Machine Learning · Computer Science 2023-12-06 Yash Chandak , Shiv Shankar , Vasilis Syrgkanis , Emma Brunskill

Bayesian Optimal Experimental Design (BOED) is a powerful tool to reduce the cost of running a sequence of experiments. When based on the Expected Information Gain (EIG), design optimization corresponds to the maximization of some…

Machine Learning · Statistics 2025-03-14 Jacopo Iollo , Christophe Heinkelé , Pierre Alliez , Florence Forbes

We propose a novel machine learning approach for inferring causal variables of a target variable from observations. Our focus is on directly inferring a set of causal factors without requiring full causal graph reconstruction, which is…

Machine Learning · Computer Science 2025-10-01 Jang-Hyun Kim , Claudia Skok Gibbs , Sangdoo Yun , Hyun Oh Song , Kyunghyun Cho

The aim of this paper is twofold. First, three theoretical principles are formalized: randomization, overrepresentation and restriction. We develop these principles and give a rationale for their use in choosing the sampling design in a…

Methodology · Statistics 2016-12-16 Yves Tillé , Matthieu Wilhelm

Clustering and dependence are common in trials. For example, in some cluster randomized trials (CRTs), pre-existing clusters are enrolled, randomized, and serve as the basis of intervention delivery. Such CRTs are "fully clustered":…

Background: When planning a cluster randomized trial, evaluators often have access to an enumerated cohort representing the target population of clusters. Practicalities of conducting the trial, such as the need to oversample clusters with…

Methodology · Statistics 2024-09-19 Sarah E. Robertson , Jon A. Steingrimsson , Issa J. Dahabreh

This paper presents the foundations of a computer oriented approach for preparing a list of random treatment assignments to be adopted in randomised controlled trials. Software is presented which can be applied in the earliest stage of…

Applications · Statistics 2015-02-12 N. S. Santos-Magalhaes , H. M. de Oliveira , A. J. Alves

We study the design of experiments with multiple treatment levels, a setting common in clinical trials and online A/B/n testing. Unlike single-treatment studies, practical analyses of multi-treatment experiments typically first select a…

Methodology · Statistics 2025-10-07 Jiachen Xu , Jian Qian , Zijun Gao

This paper is based on the observation that, during Covid-19 epidemic, the choice of which individuals should be tested has an important impact on the effectiveness of selective confinement measures. This decision problem is closely related…

Physics and Society · Physics 2020-12-24 Matthias Pezzutto , Nicolas Bono Rossello , Luca Schenato , Emanuele Garone

High-dimensional compositional data arise naturally in many applications such as metagenomic data analysis. The observed data lie in a high-dimensional simplex, and conventional statistical methods often fail to produce sensible results due…

Methodology · Statistics 2016-01-19 Yuanpei Cao , Wei Lin , Hongzhe Li

We consider a setting in which we have a treatment and a large number of covariates for a set of observations, and wish to model their relationship with an outcome of interest. We propose a simple method for modeling interactions between…

Methodology · Statistics 2012-12-14 Lu Tian , Ash Alizadeh , Andrew Gentles , Robert Tibshirani

Purpose: Machine learning is broadly used for clinical data analysis. Before training a model, a machine learning algorithm must be selected. Also, the values of one or more model parameters termed hyper-parameters must be set. Selecting…

Machine Learning · Computer Science 2018-12-10 Xueqiang Zeng , Gang Luo