中文
相关论文

相关论文: Why stratification may hurt, & how much

200 篇论文

From social networks to P2P systems, network sampling arises in many settings. We present a detailed study on the nature of biases in network sampling strategies to shed light on how best to sample from networks. We investigate connections…

社会与信息网络 · 计算机科学 2011-09-20 Arun S. Maiya , Tanya Y. Berger-Wolf

Numerous papers ask how difficult it is to cluster data. We suggest that the more relevant and interesting question is how difficult it is to cluster data sets {\em that can be clustered well}. More generally, despite the ubiquity and the…

机器学习 · 计算机科学 2012-05-23 Amit Daniely , Nati Linial , Michael Saks

The statistical efficiency of randomized clinical trials can be improved by incorporating information from baseline covariates (i.e., pre-treatment patient characteristics). This can be done in the design stage using stratified (permutated…

统计方法学 · 统计学 2025-02-04 Zhiwei Zhang

This paper develops a unified framework for partial identification and inference in stratified experiments with attrition, accommodating both equal and heterogeneous treatment shares across strata. For equal-share designs, we apply recent…

计量经济学 · 经济学 2026-01-21 Bruno Ferman , Davi Siqueira , Vitor Possebom

Temporal networks have been increasingly used to model a diversity of systems that evolve in time; for example human contact structures over which dynamic processes such as epidemics take place. A fundamental aspect of real-life networks is…

物理与社会 · 物理学 2017-11-08 Luis E C Rocha , Naoki Masuda , Petter Holme

The fragility of modern machine learning models has drawn a considerable amount of attention from both academia and the public. While immense interests were in either crafting adversarial attacks as a way to measure the robustness of neural…

机器学习 · 计算机科学 2021-03-16 Jeet Mohapatra , Ching-Yun Ko , Tsui-Wei , Weng , Sijia Liu , Pin-Yu Chen , Luca Daniel

In empirical work it is common to estimate parameters of models and report associated standard errors that account for "clustering" of units, where clusters are defined by factors such as geography. Clustering adjustments are typically…

统计理论 · 数学 2022-09-21 Alberto Abadie , Susan Athey , Guido Imbens , Jeffrey Wooldridge

Experimental work regularly finds that individual choices are not deterministically rationalized by well-defined preferences. Nonetheless, recent work shows that data collected from many individuals can be stochastically rationalized by a…

理论经济学 · 经济学 2021-10-22 Changkuk Im , John Rehbeck

We study the effectiveness of subagging, or subsample aggregating, on regression trees, a popular non-parametric method in machine learning. First, we give sufficient conditions for pointwise consistency of trees. We formalize that (i) the…

机器学习 · 统计学 2024-04-03 Christos Revelas , Otilia Boldea , Bas J. M. Werker

Design of experiments, random search, initialization of population-based methods, or sampling inside an epoch of an evolutionary algorithm use a sample drawn according to some probability distribution for approximating the location of an…

神经与进化计算 · 计算机科学 2020-04-27 Laurent Meunier , Carola Doerr , Jeremy Rapin , Olivier Teytaud

Accurately measuring discrimination is crucial to faithfully assessing fairness of trained machine learning (ML) models. Any bias in measuring discrimination leads to either amplification or underestimation of the existing disparity.…

机器学习 · 计算机科学 2025-03-25 Sami Zhioua , Ruta Binkyte , Ayoub Ouni , Farah Barika Ktata

Schools with the highest average student performance are often the smallest schools; localities with the highest rates of some cancers are frequently small and the effects observed in clinical trials are likely to be largest for the…

概率论 · 数学 2016-12-07 Steven N. Evans , Ronald L. Rivest , Philip B. Stark

In classification problems, the purpose of feature selection is to identify a small, highly discriminative subset of the original feature set. In many applications, the dataset may have thousands of features and only a few dozens of samples…

机器学习 · 计算机科学 2020-08-28 Ludmila I. Kuncheva , Clare E. Matthews , Álvar Arnaiz-González , Juan J. Rodríguez

Estimation of treatment effect for principal strata has been studied for more than two decades. Existing research exclusively focuses on the estimation, but there is little research on forming and testing hypotheses for principal…

统计方法学 · 统计学 2021-12-20 Yongming Qu , Stephen J. Ruberg , Junxiang Luo , Ilya Lipkovich

This paper proposes an adaptive randomization procedure for two-stage randomized controlled trials. The method uses data from a first-wave experiment in order to determine how to stratify in a second wave of the experiment, where the…

计量经济学 · 经济学 2022-07-06 Max Tabord-Meehan

Online controlled experiments, also known as A/B testing, are the digital equivalent of randomized controlled trials for estimating the impact of marketing campaigns on website visitors. Stratified sampling is a traditional technique for…

When using machine learning for automated prediction, it is important to account for fairness in the prediction. Fairness in machine learning aims to ensure that biases in the data and model inaccuracies do not lead to discriminatory…

机器学习 · 计算机科学 2024-12-10 Jan Pablo Burgard , João Vitor Pamplona

Stratification and rerandomization are two well-known methods used in randomized experiments for balancing the baseline covariates. Renowned scholars in experimental design have recommended combining these two methods; however, limited…

统计方法学 · 统计学 2021-10-27 Xinhe Wang , Tingyu Wang , Hanzhong Liu

Importance-weighting is a popular and well-researched technique for dealing with sample selection bias and covariate shift. It has desirable characteristics such as unbiasedness, consistency and low computational complexity. However,…

机器学习 · 统计学 2019-03-12 Wouter M. Kouw , Marco Loog

An evolving problem in the field of spatial and ecological statistics is that of preferential sampling, where biases may be present due to a relationship between sample data locations and a response of interest. This field of research bears…

统计方法学 · 统计学 2022-03-11 Daniel Vedensky , Paul A. Parker , Scott H. Holan