中文
相关论文

相关论文: Sample size planning for conditional counterfactua…

200 篇论文

Conformal prediction is a simple and powerful tool that can quantify uncertainty without any distributional assumptions. Many existing methods only address the average coverage guarantee, which is not ideal compared to the stronger…

机器学习 · 统计学 2023-02-21 Xing Han , Ziyang Tang , Joydeep Ghosh , Qiang Liu

Machine learning and statistics typically focus on building models that capture the vast majority of the data, possibly ignoring a small subset of data as "noise" or "outliers." By contrast, here we consider the problem of jointly…

机器学习 · 计算机科学 2016-08-19 Brendan Juba

Causal inference from observational data provides strong evidence for the best action in decision-making without performing expensive randomized trials. The effect of an action is usually not identifiable under unobserved confounding, even…

机器学习 · 计算机科学 2026-02-02 Md Musfiqur Rahman , Ziwei Jiang , Hilaf Hasson , Murat Kocaoglu

Randomized experiments can provide unbiased estimates of sample average treatment effects. However, estimates of population treatment effects can be biased when the experimental sample and the target population differ. In this case, the…

统计方法学 · 统计学 2022-11-10 Wenqi Shi , Xi Lin

Modern statistical analysis often encounters datasets with large sizes. For these datasets, conventional estimation methods can hardly be used immediately because practitioners often suffer from limited computational resources. In most…

统计方法学 · 统计学 2023-04-14 Shuyuan Wu , Xuening Zhu , Hansheng Wang

We address counterfactual analysis in empirical models of games with partially identified parameters, and multiple equilibria and/or randomized strategies, by constructing and analyzing the counterfactual predictive distribution set (CPDS).…

计量经济学 · 经济学 2024-10-17 Brendan Kline , Elie Tamer

Estimation of the complete distribution of a random variable is a useful primitive for both manual and automated decision making. This problem has received extensive attention in the i.i.d. setting, but the arbitrary data dependent setting…

机器学习 · 统计学 2023-03-01 Paul Mineiro , Steven R. Howard

A complete understanding of heterogeneous treatment effects involves characterizing the full conditional distribution of potential outcomes. To this end, we propose the Conditional Counterfactual Mean Embeddings (CCME), a framework that…

机器学习 · 统计学 2026-02-05 Thatchanon Anancharoenkij , Donlapark Ponnoprat

Randomization tests are a popular method for testing causal effects in clinical trials with finite-sample validity. In the presence of heterogeneous treatment effects, it is often of interest to select a subgroup that benefits from the…

统计方法学 · 统计学 2025-04-29 Zijun Gao

This paper describes three methods for carrying out non-asymptotic inference on partially identified parameters that are solutions to a class of optimization problems. Applications in which the optimization problems arise include estimation…

统计方法学 · 统计学 2022-12-02 Joel L. Horowitz , Sokbae Lee

A common problem in analysis of experiments or in lattice QCD simulations is fitting a parameterized model to the average over a number of samples of correlated data values. If the number of samples is not infinite, estimates of the…

高能物理 - 格点 · 物理学 2008-08-27 D. Toussaint , W. Freeman

The problem tackled in this paper is the determination of sample size for a given level and power in the context of a simple linear regression model. At a technical level, the simple linear regression model is a five-parameter model. It is…

统计方法学 · 统计学 2019-07-25 Tianyuan Guan , M. Khorshed Alam , M. Bhaskara Rao

We present convincing empirical evidence for an effective and general strategy for building accurate small models. Such models are attractive for interpretability and also find use in resource-constrained environments. The strategy is to…

机器学习 · 计算机科学 2024-04-30 Abhishek Ghose

Background: When developing a clinical prediction model using time-to-event data, previous research focuses on the sample size to minimise overfitting and precisely estimate the overall risk. However, instability of individual-level risk…

Practitioners often use data from a randomized controlled trial to learn a treatment assignment policy that can be deployed on a target population. A recurring concern in doing so is that, even if the randomized trial was well-executed…

计量经济学 · 经济学 2023-04-25 Lihua Lei , Roshni Sahoo , Stefan Wager

We analyze a compression scheme for large data sets that randomly keeps a small percentage of the components of each data sample. The benefit is that the output is a sparse matrix and therefore subsequent processing, such as PCA or K-means,…

机器学习 · 统计学 2017-02-24 Farhad Pourkamali-Anaraki , Stephen Becker

When using machine learning for imbalanced binary classification problems, it is common to subsample the majority class to create a (more) balanced training dataset. This biases the model's predictions because the model learns from data…

机器学习 · 计算机科学 2025-11-03 Nathan Phelps , Daniel J. Lizotte , Douglas G. Woolford

Hierarchically-organized data arise naturally in many psychology and neuroscience studies. As the standard assumption of independent and identically distributed samples does not hold for such data, two important problems are to accurately…

统计理论 · 数学 2018-09-03 Irene Dowding , Stefan Haufe

In a split conformal framework with $K$ classes, a calibration sample of $n$ labeled examples is observed for inference on the label of a new unlabeled example. We explore the setting where a `batch' of $m$ independent such unlabeled…

统计方法学 · 统计学 2025-03-19 Ulysse Gazin , Ruth Heller , Etienne Roquain , Aldo Solari

Causal effects are often characterized with population summaries. These might provide an incomplete picture when there are heterogeneous treatment effects across subgroups. Since the subgroup structure is typically unknown, it is more…

统计方法学 · 统计学 2026-04-07 Kwangho Kim , Jisu Kim , Edward H. Kennedy