English
Related papers

Related papers: High-dimensional estimation with missing data: Sta…

200 papers

Longitudinal or panel data can be represented as a matrix with rows indexed by units and columns indexed by time. We consider inferential questions associated with the missing data version of panel data induced by staggered adoption. We…

Statistics Theory · Mathematics 2024-07-02 Yuling Yan , Martin J. Wainwright

Identifying anomalies and contamination in datasets is important in a wide variety of settings. In this paper, we describe a new technique for estimating contamination in large, discrete valued datasets. Our approach considers the normal…

Information Theory · Computer Science 2015-06-16 Matthew L. Malloy , Scott Alfeld , Paul Barford

Matrix completion is often applied to data with entries missing not at random (MNAR). For example, consider a recommendation system where users tend to only reveal ratings for items they like. In this case, a matrix completion method that…

Machine Learning · Statistics 2019-10-30 Wei Ma , George H. Chen

Estimation problems with constrained parameter spaces arise in various settings. In many of these problems, the observations available to the statistician can be modelled as arising from the noisy realization of the image of a random linear…

Statistics Theory · Mathematics 2023-03-23 Reese Pathak , Martin J. Wainwright , Lin Xiao

Matrix completion is the study of recovering an underlying matrix from a sparse subset of noisy observations. Traditionally, it is assumed that the entries of the matrix are "missing completely at random" (MCAR), i.e., each entry is…

Econometrics · Economics 2021-10-01 Anish Agarwal , Munther Dahleh , Devavrat Shah , Dennis Shen

Classical semiparametric inference with missing outcome data is not robust to contamination of the observed data and a single observation can have arbitrarily large influence on estimation of a parameter of interest. This sensitivity is…

Methodology · Statistics 2021-03-02 Eva Cantoni , Xavier de Luna

Distributed systems have been widely used in practice to accomplish data analysis tasks of huge scales. In this work, we target on the estimation problem of generalized linear models on a distributed system with nonrandomly distributed…

Methodology · Statistics 2020-04-07 Feifei Wang , Danyang Huang , Yingqiu Zhu , Hansheng Wang

We study sparse linear regression over a network of agents, modeled as an undirected graph (with no centralized node). The estimation problem is formulated as the minimization of the sum of the local LASSO loss functions plus a quadratic…

Machine Learning · Computer Science 2023-06-23 Yao Ji , Gesualdo Scutari , Ying Sun , Harsha Honnappa

Recently, many machine learning and statistical models such as non-linear regressions, the Single Index, Multi-index, Varying Coefficient Index Models and Two-layer Neural Networks can be reduced to or be seen as a special case of a new…

Machine Learning · Computer Science 2020-10-20 Di Wang , Xiangyu Guo , Chaowen Guan , Shi Li , Jinhui Xu

As one of the most commonly seen data challenges, missing data, in particular, multiple, non-monotone missing patterns, complicates estimation and inference due to the fact that missingness mechanisms are often not missing at random, and…

Methodology · Statistics 2025-04-21 Jianing Dong , Raymond K. W. Wong , Kwun Chuen Gary Chan

During the past few decades, missing-data problems have been studied extensively, with a focus on the ignorable missing case, where the missing probability depends only on observable quantities. By contrast, research into non-ignorable…

Methodology · Statistics 2019-08-06 Yukun Liu , Pengfei Li , Jing Qin

Structural matrix-variate observations routinely arise in diverse fields such as multi-layer network analysis and brain image clustering. While data of this type have been extensively investigated with fruitful outcomes being delivered, the…

Statistics Theory · Mathematics 2022-01-25 Zhongyuan Lyu , Dong Xia

Machine learning and statistics typically focus on building models that capture the vast majority of the data, possibly ignoring a small subset of data as "noise" or "outliers." By contrast, here we consider the problem of jointly…

Machine Learning · Computer Science 2016-08-19 Brendan Juba

Deep neural networks (DNNs) have become powerful tools for modeling complex data structures through sequentially integrating simple functions in each hidden layer. In survival analysis, recent advances of DNNs primarily focus on enhancing…

Machine Learning · Statistics 2025-03-26 Changhui Yuan , Shishun Zhao , Shuwei Li , Xinyuan Song , Zhao Chen

In some multivariate problems with missing data, pairs of variables exist that are never observed together. For example, some modern biological tools can produce data of this form. As a result of this structure, the covariance matrix is…

Methodology · Statistics 2013-08-13 Max Grazier G'Sell , Shai S. Shen-Orr , Robert Tibshirani

In the analysis of observational data in social sciences and businesses, it is difficult to obtain a "(quasi) single-source dataset" in which the variables of interest are simultaneously observed. Instead, multiple-source datasets are…

Methodology · Statistics 2021-09-02 Masaki Mitsuhiro , Takahiro Hoshino

Many statistical estimators are defined as the fixed point of a data-dependent operator, with estimators based on minimizing a cost function being an important special case. The limiting performance of such estimators depends on the…

Machine Learning · Computer Science 2022-03-22 Nhat Ho , Koulik Khamaru , Raaz Dwivedi , Martin J. Wainwright , Michael I. Jordan , Bin Yu

Missing data is an universal problem in statistics. We develop a unified framework for estimating parameters defined by general estimating equations under a missing-at-random (MAR) mechanism, based on generalized entropy calibration…

Methodology · Statistics 2026-03-31 Mst Moushumi Pervin , Hengfang Wang , Jae Kwang Kim

Missing covariates are not uncommon in capture-recapture studies. When covariate information is missing at random in capture-recapture data, an empirical full likelihood method has been demonstrated to outperform…

Methodology · Statistics 2025-07-15 Yang Liu , Yukun Liu , Pengfei Li , Riquan Zhang

We develop new methods to integrate experimental and observational data in causal inference. While randomized controlled trials offer strong internal validity, they are often costly and therefore limited in sample size. Observational data,…

Econometrics · Economics 2025-11-04 Xuelin Yang , Licong Lin , Susan Athey , Michael I. Jordan , Guido W. Imbens