中文
相关论文

相关论文: Endogenous post-stratification in surveys: classif…

200 篇论文

The problems that exist in implementing a sampling design for socio-economic surveys in remote areas in Indonesia are high cost of the survey, low response rate, and less accurate. Therefore, the sampling design needs to be developed, one…

统计方法学 · 统计学 2022-12-07 Adhi Kurniawan , Atika Nashirah Hasyyati

Online reinforcement learning and other adaptive sampling algorithms are increasingly used in digital intervention experiments to optimize treatment delivery for users over time. In this work, we focus on longitudinal user data collected by…

机器学习 · 计算机科学 2023-04-20 Kelly W. Zhang , Lucas Janson , Susan A. Murphy

Multilevel regression and poststratification (MRP) is a flexible modeling technique that has been used in a broad range of small-area estimation problems. Traditionally, MRP studies have been focused on non-causal settings, where estimating…

统计方法学 · 统计学 2022-01-24 Yuxiang Gao , Lauren Kennedy , Daniel Simpson

As predictive algorithms grow in popularity, using the same dataset to both train and test a new model has become routine across research, policy, and industry. Sample-splitting attains valid inference on model properties by using separate…

计量经济学 · 经济学 2025-11-27 Bruno Fava

This paper develops a unified framework for partial identification and inference in stratified experiments with attrition, accommodating both equal and heterogeneous treatment shares across strata. For equal-share designs, we apply recent…

计量经济学 · 经济学 2026-01-21 Bruno Ferman , Davi Siqueira , Vitor Possebom

Model performance evaluation is a critical and expensive task in machine learning and computer vision. Without clear guidelines, practitioners often estimate model accuracy using a one-time completely random selection of the data. However,…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Riccardo Fogliato , Pratik Patil , Mathew Monfort , Pietro Perona

Class imbalance in real-world data poses a common bottleneck for machine learning tasks, since achieving good generalization on under-represented examples is often challenging. Mitigation strategies, such as under or oversampling the data…

无序系统与神经网络 · 物理学 2025-02-03 Emanuele Loffredo , Mauro Pastore , Simona Cocco , Rémi Monasson

Nonparametric cointegrating regression models have been extensively used in financial markets, stock prices, heavy traffic, climate data sets, and energy markets. Models with parametric regression functions can be more appealing in practice…

统计方法学 · 统计学 2023-12-27 Sepideh Mosaferi , Mark S. Kaiser , Daniel J. Nordman

A balanced sampling design should always be the adopted strategies if auxiliary information is available. Besides, integrating a stratified structure of the population in the sampling process can considerably reduce the variance of the…

统计方法学 · 统计学 2022-06-03 Raphaël Jauslin , Esther Eustache , Yves Tillé

Analysis of sample survey data often requires adjustments to account for missing data in the outcome variables of principal interest. Standard adjustment methods based on item imputation or on propensity weighting factors rely heavily on…

统计方法学 · 统计学 2016-03-08 Wei-Yin Loh , John Eltinge , MoonJung Cho , Yuanzhi Li

The need for rigorous and timely health and demographic summaries has provided the impetus for an explosion in geographic studies, with a common approach being the production of pixel-level maps, particularly in low and middle income…

统计方法学 · 统计学 2019-10-16 John Paige , Geir-Arne Fuglstad , Andrea Riebler , Jon Wakefield

We propose a simple, statistically principled, and theoretically justified method to improve supervised learning when the training set is not representative, a situation known as covariate shift. We build upon a well-established methodology…

机器学习 · 统计学 2025-03-12 Maximilian Autenrieth , David A. van Dyk , Roberto Trotta , David C. Stenning

When modelling data where the response is dichotomous and highly imbalanced, response-based sampling where a subset of the majority class is retained (i.e., undersampling) is often used to create more balanced training datasets prior to…

统计方法学 · 统计学 2024-12-06 Nathan Phelps , Daniel J. Lizotte , Douglas G. Woolford

Systematic sampling is often used to select plot locations for forest inventory estimation. However, it is not possible to derive a design-unbiased variance estimator for a systematic sample using one random start. As a result, many forest…

应用统计 · 统计学 2018-10-22 Chad Babcock , Andrew O. Finley , Timothy G. Gregoire , Hans-Erik Andersen

Practitioners are interested in not only the average causal effect of the treatment on the outcome but also the underlying causal mechanism in the presence of an intermediate variable between the treatment and outcome. However, in many…

统计方法学 · 统计学 2016-02-04 Peng Ding , Jiannan Lu

This paper proposes several tests of restricted specification in nonparametric instrumental regression. Based on series estimators, test statistics are established that allow for tests of the general model against a parametric or…

计量经济学 · 经济学 2019-09-24 Christoph Breunig

Probabilistic classifiers output a probability distribution on target classes rather than just a class prediction. Besides providing a clear separation of prediction and decision making, the main advantage of probabilistic models is their…

机器学习 · 计算机科学 2019-02-20 Juozas Vaicenavicius , David Widmann , Carl Andersson , Fredrik Lindsten , Jacob Roll , Thomas B. Schön

Several techniques exist to assess and reduce nonresponse bias, including propensity models, calibration methods, or post-stratification. These approaches can only be applied after the data collection, and assume reliable information…

统计方法学 · 统计学 2020-05-26 Blanka Szeitl , Tamás Rudas

For consistency (even oracle properties) of estimation and model prediction, almost all existing methods of variable/feature selection critically depend on sparsity of models. However, for ``large $p$ and small $n$" models sparsity…

统计方法学 · 统计学 2010-08-10 Lu Lin , Lixing Zhu , Yujie Gai

In classification problems, sampling bias between training data and testing data is critical to the ranking performance of classification scores. Such bias can be both unintentionally introduced by data collection and intentionally…

统计方法学 · 统计学 2017-11-02 Chandler Zuo