English
Related papers

Related papers: Dealing with Logs and Zeros in Regression Models

200 papers

Forecast-then-optimize is a widely-used framework for decision-making problems in power systems. Traditionally, statistical losses have been employed to train forecasting models, but recent research demonstrated that improved decision…

Systems and Control · Electrical Eng. & Systems 2023-12-22 Haipeng Zhang , Ran Li , Mingyang Sun , Teng Fei

Under the Neyman causal model, it is well-known that OLS with treatment-by-covariate interactions cannot harm asymptotic precision of estimated treatment effects in completely randomized experiments. But do such guarantees extend to…

Statistics Theory · Mathematics 2018-03-19 Joel A. Middleton

This article introduces a leave-one-out regression adjustment (LOORA) for estimating average treatment effects in randomized controlled trials. In finite samples, LOORA removes the bias of conventional regression adjustment and yields exact…

Econometrics · Economics 2026-05-08 Alberto Abadie , Mehrdad Ghadiri , Ali Jadbabaie , Mahyar JafariNodeh

AI for IT Operations (AIOps) is a powerful platform that Site Reliability Engineers (SREs) use to automate and streamline operational workflows with minimal human intervention. Automated log analysis is a critical task in AIOps as it…

Computation and Language · Computer Science 2023-08-23 Pranjal Gupta , Harshit Kumar , Debanjana Kar , Karan Bhukar , Pooja Aggarwal , Prateeti Mohapatra

We study the asymptotic properties of the adaptive Lasso in cointegration regressions in the case where all covariates are weakly exogenous. We assume the number of candidate I(1) variables is sub-linear with respect to the sample size (but…

Methodology · Statistics 2011-10-11 Eduardo F. Mendes

This paper proposes a method for estimating multiple change points in panel data models with unobserved individual effects via ordinary least-squares (OLS). Typically, in this setting, the OLS slope estimators are inconsistent due to the…

Econometrics · Economics 2018-08-10 Otilia Boldea , Bettina Drepper , Zhuojiong Gan

We provide a principled way for investigators to analyze randomized experiments when the number of covariates is large. Investigators often use linear multivariate regression to analyze randomized experiments instead of simply reporting the…

Statistics Theory · Mathematics 2022-06-08 Adam Bloniarz , Hanzhong Liu , Cun-Hui Zhang , Jasjeet Sekhon , Bin Yu

This paper studies estimation of linear panel regression models with heterogeneous coefficients, when both the regressors and the residual contain a possibly common, latent, factor structure. Our theory is (nearly) efficient, because based…

Econometrics · Economics 2019-03-01 Marco Avarucci , Paolo Zaffaroni

Personalized medicine aims to tailor treatments to individual patients, especially when people respond heterogeneously to therapies. A key objective is to learn individualized treatment rules that recommend optimal treatments from patient…

Methodology · Statistics 2026-03-12 Zhu Wang

This paper considers a model with general regressors and unobservable factors. An estimator based on iterated principal components is proposed, which is shown to be not only asymptotically normal and oracle efficient, but under certain…

Econometrics · Economics 2025-04-23 Bin Peng , Liangjun Su , Joakim Westerlund , Yanrong Yang

We present a comprehensive framework for applying rigorous statistical techniques from econometrics to analyze and improve machine learning systems. We introduce key statistical methods such as Ordinary Least Squares (OLS) regression,…

Machine Learning · Computer Science 2024-10-03 Michaël Soumm

This letter deals with a very simple issue: if we have grouped data with a binary dependent variable and want to include fixed effects (group specific intercepts) in the specification, is Ordinary Least Squares (OLS) in any way superior to…

Applications · Statistics 2018-10-30 Nathaniel Beck

The fate of scientific hypotheses often relies on the ability of a computational model to explain the data, quantified in modern statistical approaches by the likelihood function. The log-likelihood is the key element for parameter…

Machine Learning · Computer Science 2021-01-27 Bas van Opheusden , Luigi Acerbi , Wei Ji Ma

In regression models involving economic variables such as income, log transformation is typically taken to achieve approximate normality and stabilize the variance. However, often the interest is predicting individual values or means of the…

Statistics Theory · Mathematics 2016-10-25 Nirian Martin , Isabel Molina

This paper studies linear time series regressions with many regressors. Weak exogeneity is the most used identifying assumption in time series. Weak exogeneity requires the structural error to have zero conditional expectation given the…

Econometrics · Economics 2024-01-18 Anna Mikusheva , Mikkel Sølvsten

This study investigated the problem posed by using ordinary least squares (OLS) to estimate parameters of simple linear regression under a specific context of special relativity, where an independent variable is restricted to an open…

Other Statistics · Statistics 2020-06-01 Si Hyung Joo

This study proposes an econometric framework to interpret and empirically decompose the difference between IV and OLS estimates given by a linear regression model when the true causal effects of the treatment are nonlinear in treatment…

Econometrics · Economics 2022-06-27 Shoya Ishimaru

In this paper, we propose an adaptive group lasso procedure to efficiently estimate structural breaks in cointegrating regressions. It is well-known that the group lasso estimator is not simultaneously estimation consistent and model…

Econometrics · Economics 2021-04-21 Karsten Schweikert

Extraordinary amounts of data are being produced in many branches of science. Proven statistical methods are no longer applicable with extraordinary large data sets due to computational limitations. A critical step in big data analysis is…

Methodology · Statistics 2019-06-27 HaiYing Wang , Min Yang , John Stufken

The dramatic growth of big datasets presents a new challenge to data storage and analysis. Data reduction, or subsampling, that extracts useful information from datasets is a crucial step in big data analysis. We propose an orthogonal…

Methodology · Statistics 2021-06-01 Lin Wang , Jake Elmstedt , Weng Kee Wong , Hongquan Xu