English
Related papers

Related papers: Insights from Machine Learning for Evaluating Prod…

200 papers

We study a linear high-dimensional regression model in a semi-supervised setting, where for many observations only the vector of covariates $X$ is given with no response $Y$. We do not make any sparsity assumptions on the vector of…

Statistics Theory · Mathematics 2021-09-03 Ilan Livne , David Azriel , Yair Goldberg

CoWrangler is a data-wrangling recommender system designed to streamline data processing tasks. Recognizing that data processing is often time-consuming and complex for novice users, we aim to simplify the decision-making process regarding…

Databases · Computer Science 2024-09-18 Yuqing Wang , Anna Fariha

Feature selection problems have been extensively studied for linear estimation, for instance, Lasso, but less emphasis has been placed on feature selection for non-linear functions. In this study, we propose a method for feature selection…

Machine Learning · Computer Science 2020-07-28 Yutaro Yamada , Ofir Lindenbaum , Sahand Negahban , Yuval Kluger

In linear regression we wish to estimate the optimum linear least squares predictor for a distribution over $d$-dimensional input points and real-valued responses, based on a small sample. Under standard random design analysis, where the…

Machine Learning · Statistics 2022-06-08 Michał Dereziński , Manfred K. Warmuth , Daniel Hsu

We propose a new estimator for the high-dimensional linear regression model with observation error in the design where the number of coefficients is potentially larger than the sample size. The main novelty of our procedure is that the…

Methodology · Statistics 2019-09-09 Alexandre Belloni , Abhishek Kaul , Mathieu Rosenbaum

Purpose: Trading on electricity markets occurs such that the price settlement takes place before delivery, often day-ahead. In practice, these prices are highly volatile as they largely depend upon a range of variables such as electricity…

Applications · Statistics 2020-05-19 Christof Naumzik , Stefan Feuerriegel

We consider the estimation of a structural function which models a non-parametric relationship between a response and an endogenous regressor given an instrument in presence of dependence in the data generating process. Assuming an…

Statistics Theory · Mathematics 2016-04-08 Nicolas Asin , Jan Johannes

We determine the expected error by smoothing the data locally. Then we optimize the shape of the kernel smoother to minimize the error. Because the optimal estimator depends on the unknown function, our scheme automatically adjusts to the…

Methodology · Statistics 2019-11-19 Kurt S. Riedel , A. Sidorenko

We consider a longitudinal data structure consisting of baseline covariates, time-varying treatment variables, intermediate time-dependent covariates, and a possibly time dependent outcome. Previous studies have shown that estimating the…

Statistics Theory · Mathematics 2018-10-09 Linh Tran , Maya Petersen , Joshua Schwab , Mark J van der Laan

The popularity of online surveys has increased the prominence of using weights that capture units' probabilities of inclusion for claims of representativeness. Yet, much uncertainty remains regarding how these weights should be employed in…

Methodology · Statistics 2017-08-16 Luke W. Miratrix , Jasjeet S. Sekhon , Alexander G. Theodoridis , Luis F. Campos

Industrial applications of machine learning face unique challenges due to the nature of raw industry data. Preprocessing and preparing raw industrial data for machine learning applications is a demanding task that often takes more time and…

Machine Learning · Computer Science 2021-09-09 Philipp Fleck , Manfred Kügel , Michael Kommenda

In this paper, we propose the application of shrinkage strategies to estimate coefficients in the Bell regression models when prior information about the coefficients is available. The Bell regression models are well-suited for modeling…

Statistics Theory · Mathematics 2024-01-03 Solmaz Seifollahi , Hossein Bevrani , Zakariya Yahya Algamal

Estimating mutual information (MI) is a fundamental task in data science and machine learning. Existing estimators mainly rely on either highly flexible models (e.g., neural networks), which require large amounts of data, or overly…

Machine Learning · Computer Science 2025-10-27 Yanzhi Chen , Zijing Ou , Adrian Weller , Michael U. Gutmann

We consider the problem of estimating how well a model class is capable of fitting a distribution of labeled data. We show that it is often possible to accurately estimate this "learnability" even when given an amount of data that is too…

Machine Learning · Computer Science 2019-03-26 Weihao Kong , Gregory Valiant

Selection bias is a serious potential problem for inference about relationships of scientific interest based on samples without well-defined probability sampling mechanisms. Motivated by the potential for selection bias in (a) estimated…

Many causal and structural effects depend on regressions. Examples include policy effects, average derivatives, regression decompositions, average treatment effects, causal mediation, and parameters of economic structural models. The…

Statistics Theory · Mathematics 2022-10-25 Victor Chernozhukov , Whitney K Newey , Rahul Singh

Assessing sensitivity to unmeasured confounding is an important step in observational studies, which typically estimate effects under the assumption that all confounders are measured. In this paper, we develop a sensitivity analysis…

Methodology · Statistics 2023-09-04 Dan Soriano , Eli Ben-Michael , Peter J. Bickel , Avi Feller , Samuel D. Pimentel

In the age of big data, nonprobability surveys are becoming increasingly abundant. Data integration techniques involving both probability and nonprobability surveys are being extensively used for providing improved estimates for finite…

Applications · Statistics 2025-10-17 Aditi Sen , Partha Lahiri

In practice, data often contain discrete variables. But most of the popular nonparametric estimation methods have been developed in a purely continuous framework. A common trick among practitioners is to make discrete variables continuous…

Methodology · Statistics 2018-01-08 Thomas Nagler

We develop pre-trained estimators for structural econometric models. The estimator uses a neural net to recognize the structural model's parameter from data patterns. Once trained, the estimator can be shared and applied to different…

Econometrics · Economics 2025-12-01 Yanhao 'Max' Wei , Zhenling Jiang