English
Related papers

Related papers: Dyadic Regression with Sample Selection

200 papers

Central banks rely on density forecasts from professional surveys to assess inflation risks and communicate uncertainty. A central challenge in using these surveys is irregular participation: forecasters enter and exit, skip rounds, and…

Applications · Statistics 2026-02-06 Matthew C. Johnson , Matteo Luciani , Minzhengxiong Zhang , Kenichiro McAlinn

This paper investigates the predictive performance of model averaging in high-dimensional linear regression where the number of regressors is comparable to the sample size. We demonstrate that the double descent trajectory manifests within…

Methodology · Statistics 2026-05-14 Ke Chen , Dandan Jiang , Xinyu Zhang

The paper considers variable selection in linear regression models where the number of covariates is possibly much larger than the number of observations. High dimensionality of the data brings in many complications, such as (possibly…

Methodology · Statistics 2016-11-29 Haeran Cho , Piotr Fryzlewicz

We consider the problem of heteroscedastic linear regression, where, given $n$ samples $(\mathbf{x}_i, y_i)$ from $y_i = \langle \mathbf{w}^{*}, \mathbf{x}_i \rangle + \epsilon_i \cdot \langle \mathbf{f}^{*}, \mathbf{x}_i \rangle$ with…

Machine Learning · Statistics 2023-07-04 Dheeraj Baby , Aniket Das , Dheeraj Nagaraj , Praneeth Netrapalli

An adaptive nonparametric estimation procedure is constructed for the estimation problem of heteroscedastic regression when the noise variance depends on the unknown regression. A non-asymptotic upper bound for a quadratic risk (an oracle…

Statistics Theory · Mathematics 2008-12-18 Leonid Galtchouk , Serguey Pergamenshchikov

In this paper, we consider several efficient data structures for the problem of sampling from a dynamically changing discrete probability distribution, where some prior information is known on the distribution of the rates, in particular…

Computational Engineering, Finance, and Science · Computer Science 2021-10-13 Federico D'Ambrosio , Hans L. Bodlaender , Gerard T. Barkema

Distributed statistical learning problems arise commonly when dealing with large datasets. In this setup, datasets are partitioned over machines, which compute locally, and communicate short messages. Communication is often the bottleneck.…

Statistics Theory · Mathematics 2022-10-25 Edgar Dobriban , Yue Sheng

The use of machine learning methods for predictive purposes has increased dramatically over the past two decades, but uncertainty quantification for predictive comparisons remains elusive. This paper addresses this gap by extending the…

Econometrics · Economics 2025-05-09 Juan Carlos Escanciano , Ricardo Parra

This paper introduces a quantile regression estimator for panel data models with individual heterogeneity and attrition. The method is motivated by the fact that attrition bias is often encountered in Big Data applications. For example,…

Econometrics · Economics 2018-08-13 Matthew Harding , Carlos Lamarche

Collected data, which is used for analysis or prediction tasks, often have a hierarchical structure, for example, data from various people performing the same task. Modeling the data's structure can improve the reliability of the derived…

Applications · Statistics 2018-11-12 Dennis Becker

The machine learning random Fourier feature method for data in high dimension is computationally and theoretically attractive since the optimization is based on a convex standard least squares problem and independent sampling of Fourier…

Numerical Analysis · Mathematics 2026-05-19 Xin Huang , Aku Kammonen , Anamika Pandey , Mattias Sandberg , Erik von Schwerin , Anders Szepessy , Raúl Tempone

The finite sensitivity of instruments or detection methods means that data sets in many areas of astronomy, for example cosmological or exoplanet surveys, are necessarily systematically incomplete. Such data sets, where the population being…

Instrumentation and Methods for Astrophysics · Physics 2020-10-14 Adam B. Mantz

Given $m$ $d$-dimensional responsors and $n$ $d$-dimensional predictors, sparse regression finds at most $k$ predictors for each responsor for linear approximation, $1\leq k \leq d-1$. The key problem in sparse regression is subset…

Machine Learning · Computer Science 2020-11-25 Jianji Wang , Qi Liu , Shupei Zhang , Nanning Zheng , Fei-Yue Wang

This paper studies the case of possibly high-dimensional covariates in the regression discontinuity design (RDD) analysis. In particular, we propose estimation and inference methods for the RDD models with covariate selection which perform…

Econometrics · Economics 2026-01-21 Yoichi Arai , Taisuke Otsu , Myung Hwan Seo

A general asymptotic theory is given for the panel data AR(1) model with time series independent in different cross sections. The theory covers the cases of stationary process, nearly non-stationary process, unit root process, mildly…

Applications · Statistics 2016-11-15 Jianfei Shen , Tianxiao Pang

Subject selection plays a critical role in experimental studies, especially ones with human subjects. Anecdotal evidence suggests that many such studies, done at or near university campus settings suffer from selection bias, i.e., the…

Machine Learning · Computer Science 2020-12-21 Tahereh Arabghalizi , Alexandros Labrinidis

The concept of biased data is well known and its practical applications range from social sciences and biology to economics and quality control. These observations arise when a sampling procedure chooses an observation with probability that…

Statistics Theory · Mathematics 2007-06-13 Sam Efromovich

Astronomers are often confronted with funky populations and distributions of objects: brighter objects are more likely to be detected; targets are selected based on colour cuts; imperfect classification yields impure samples. Failing to…

Cosmology and Nongalactic Astrophysics · Physics 2017-06-21 Samuel R. Hinton , Alex Kim , Tamara M. Davis

Spatial autocorrelation in regression models can lead to downward biased standard errors and thus incorrect inference. The most common correction in applied economics is the spatial heteroskedasticity and autocorrelation consistent (HAC)…

Econometrics · Economics 2026-03-05 Alexander Lehner

The paper deals with asymptotic properties of the adaptive procedure proposed in the author paper, 2007, for estimating a unknown nonparametric regression. We prove that this procedure is asymptotically efficient for a quadratic risk, i.e.…

Statistics Theory · Mathematics 2008-10-08 Leonid Galtchouk , Serguey Pergamenshchikov