English
Related papers

Related papers: (Ab)Using Regression for Data Adjustment

200 papers

Large-scale data are often characterized by some degree of inhomogeneity as data are either recorded in different time regimes or taken from multiple sources. We look at regression models and the effect of randomly changing coefficients,…

Methodology · Statistics 2016-08-11 Nicolai Meinshausen , Peter Bühlmann

Empirical researchers often use diagnostic checks to assess the plausibility of their modeling assumptions, such as testing for covariate balance in RCTs, pre-trends in event studies, or instrument validity in IV designs. While these checks…

Econometrics · Economics 2026-04-21 Reca Sarfati , Vod Vilfort

I study identification, estimation and inference for spillover effects in experiments where units' outcomes may depend on the treatment assignments of other units within a group. I show that the commonly-used reduced-form linear-in-means…

Econometrics · Economics 2022-01-21 Gonzalo Vazquez-Bare

Heterogeneous but complementary sources of data provide an unprecedented opportunity for developing accurate statistical models of systems. Although the existing methods have shown promising results, they are mostly applicable to situations…

Applications · Statistics 2020-08-18 Feng Wang , Mostafa Reisi Gahrooei , Zhen Zhong , Tao Tang , Jianjun Shi

The advent of modern data collection and processing techniques has seen the size, scale, and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the…

Applications · Statistics 2024-07-30 Zeyi Wang , Eric Bridgeford , Shangsi Wang , Joshua T. Vogelstein , Brian Caffo

Traditional metrics like accuracy, F1-score, and precision are frequently used to evaluate machine learning models, however they may not be sufficient for evaluating performance on tiny, unbalanced, or high-dimensional datasets. A…

Machine Learning · Computer Science 2024-12-11 Serzhan Ossenov

In this paper we present tools for applied researchers that re-purpose off-the-shelf methods from the computer-science field of machine learning to create a "discovery engine" for data from randomized controlled trials (RCTs). The applied…

Machine Learning · Statistics 2019-05-13 Jens Ludwig , Sendhil Mullainathan , Jann Spiess

To develop software with optimal performance, even small performance changes need to be identified. Identifying performance changes is challenging since the performance of software is influenced by non-deterministic factors. Therefore, not…

Software Engineering · Computer Science 2023-03-28 David Georg Reichelt , Stefan Kühne , Wilhelm Hasselbring

Estimation of a conditional mean (linking a set of features to an outcome of interest) is a fundamental statistical task. While there is an appeal to flexible nonparametric procedures, effective estimation in many classical nonparametric…

Methodology · Statistics 2022-06-08 Tianyu Zhang , Noah Simon

We consider regression models with parametric (linear or nonlinear) regression function and allow responses to be ``missing at random.'' We assume that the errors have mean zero and are independent of the covariates. In order to estimate…

Statistics Theory · Mathematics 2009-08-24 Ursula U. Müller

We consider the model $Z_i=X_i+\varepsilon_i$, for i.i.d. $X_i$'s and $\varepsilon_i$'s and independent sequences $(X_i)_{i\in{\mathbb{N}}}$ and $(\varepsilon_i)_{i\in{\mathbb{N}}}$. The density $f_{\varepsilon}$ of $\varepsilon_1$ is…

Statistics Theory · Mathematics 2009-02-10 C. Butucea , F. Comte

We consider the problem of robustly fitting a model to data that includes outliers by formulating a percentile optimization problem. This problem is non-smooth and non-convex, hence hard to solve. We derive properties that the minimizers of…

Signal Processing · Electrical Eng. & Systems 2024-05-16 João Domingos , João Xavier

Most fair regression algorithms mitigate bias towards sensitive sub populations and therefore improve fairness at group level. In this paper, we investigate the impact of such implementation of fair regression on the individual. More…

Machine Learning · Computer Science 2021-04-12 Boris Ruf , Marcin Detyniecki

We investigate how to improve efficiency using regression adjustments with covariates in covariate-adaptive randomizations (CARs) with imperfect subject compliance. Our regression-adjusted estimators, which are based on the doubly robust…

Econometrics · Economics 2023-06-19 Liang Jiang , Oliver B. Linton , Haihan Tang , Yichong Zhang

Machine learning models trained on real-world data may inadvertently make biased predictions that negatively impact marginalized communities. Reweighting, which assigns a weight to each data point used during model training, can mitigate…

Machine Learning · Computer Science 2026-03-20 Anil K. Saini , Jose Guadalupe Hernandez , Emily F. Wong , Debanshi Misra , Tiffani J. Bright , Jason H. Moore

Long memory in the sense of slowly decaying autocorrelations is a stylized fact in many time series from economics and finance. The fractionally integrated process is the workhorse model for the analysis of these time series. Nevertheless,…

Econometrics · Economics 2023-09-22 Uwe Hassler , Marc-Oliver Pohle

Variance function estimation in nonparametric regression is considered and the minimax rate of convergence is derived. We are particularly interested in the effect of the unknown mean on the estimation of the variance function. Our results…

Statistics Theory · Mathematics 2008-12-18 Lie Wang , Lawrence D. Brown , T. Tony Cai , Michael Levine

Density regression provides a flexible strategy for modeling the distribution of a response variable $Y$ given predictors $\mathbf{X}=(X_1,\ldots,X_p)$ by letting that the conditional density of $Y$ given $\mathbf{X}$ as a completely…

Statistics Theory · Mathematics 2016-01-07 Weining Shen , Subhashis Ghosal

In this paper, we introduce a novel method to generate interpretable regression function estimators. The idea is based on called data-dependent coverings. The aim is to extract from the data a covering of the feature space instead of a…

Statistics Theory · Mathematics 2021-01-27 Vincent Margot , Jean-Patrick Baudry , Frédéric Guilloux , Olivier Wintenberger

To increase statistical efficiency in a randomized experiment, researchers often use stratification (i.e., blocking) in the design stage. However, conventional practices of stratification fail to exploit valuable information about the…

Methodology · Statistics 2025-10-28 Zikai Li
‹ Prev 1 8 9 10 Next ›