English
Related papers

Related papers: Robust and flexible inference for the covariate-sp…

200 papers

We introduce a flexible framework for high-dimensional matrix estimation to incorporate side information for both rows and columns. Existing approaches, such as inductive matrix completion, often impose restrictive structure-for example, an…

Methodology · Statistics 2026-03-27 Anish Agarwal , Jungjun Choi , Ming Yuan

In the fight against hard-to-treat diseases such as cancer, it is often difficult to discover new treatments that benefit all subjects. For regulatory agency approval, it is more practical to identify subgroups of subjects for whom the…

Methodology · Statistics 2014-10-09 Wei-Yin Loh , Xu He , Michael Man

Flexible estimation of the mean outcome under a treatment regimen (i.e., value function) is the key step toward personalized medicine. We define our target parameter as a conditional value function given a set of baseline covariates which…

Statistics Theory · Mathematics 2023-09-29 Ashkan Ertefaie , Luke Duttweiler , Brent A. Johnson , Mark J. van der Laan

With the advance of high-throughput genotyping and sequencing technologies, it becomes feasible to comprehensive evaluate the role of massive genetic predictors in disease prediction. There exists, therefore, a critical need for developing…

Methodology · Statistics 2025-04-02 Changshuai Wei , Ming Li , Yalu Wen , Chengyin Ye , Qing Lu

Empirical regression discontinuity (RD) studies often include covariates in their specifications to increase the precision of their estimates. In this paper, we propose a novel class of estimators that use such covariate information more…

Econometrics · Economics 2025-04-28 Claudia Noack , Tomasz Olma , Christoph Rothe

Deploying clinical prediction models across healthcare systems often fails when key training covariates are unavailable at deployment and labeled outcomes are limited in the target domain. For example, high-performing models for…

Network experiments are powerful tools for studying spillover effects, which avoid endogeneity by randomly assigning treatments to units over networks. However, it is non-trivial to analyze network experiments properly without imposing…

Econometrics · Economics 2025-06-09 Mengsi Gao , Peng Ding

Multivariate linear regression is a fundamental statistical task, but classical estimators such as ordinary least squares are highly sensitive to outliers. These may occur as casewise outliers that affect entire observations, or as outlying…

Methodology · Statistics 2026-05-11 Fabio Centofanti , Mia Hubert , Peter J. Rousseeuw

Methods for the evaluation of the predictive accuracy of biomarkers with respect to survival outcomes subject to right censoring have been discussed extensively in the literature. In cancer and other diseases, survival outcomes are commonly…

Methodology · Statistics 2018-06-06 Yuan Wu , Xiaofei Wang , Jiaxing Lin , Beilin Jia , Kouros Owzar

We consider inference in linear regression models that is robust to heteroskedasticity and the presence of many control variables. When the number of control variables increases at the same rate as the sample size the usual…

Statistics Theory · Mathematics 2020-09-29 Koen Jochmans

The reliability of machine learning systems critically assumes that the associations between features and labels remain similar between training and test distributions. However, unmeasured variables, such as confounders, break this…

Machine Learning · Computer Science 2020-08-17 Megha Srivastava , Tatsunori Hashimoto , Percy Liang

For high-dimensional classification, it is well known that naively performing the Fisher discriminant rule leads to poor results due to diverging spectra and noise accumulation. Therefore, researchers proposed independence rules to…

Machine Learning · Statistics 2011-11-10 Jianqing Fan , Yang Feng , Xin Tong

The ability to generalize experimental results from randomized control trials (RCTs) across locations is crucial for informing policy decisions in targeted regions. Such generalization is often hindered by the lack of identifiability due to…

Econometrics · Economics 2021-12-10 Xinkun Nie , Guido Imbens , Stefan Wager

It is common to conduct causal inference in matched observational studies by proceeding as though treatment assignments within matched sets are assigned uniformly at random and using this distribution as the basis for inference. This…

Methodology · Statistics 2023-11-14 Samuel D. Pimentel , Yaxuan Huang

Empirical studies using Regression Discontinuity (RD) designs often explore heterogeneous treatment effects based on pretreatment covariates, even though no formal statistical methods exist for such analyses. This has led to the widespread…

In binary classification applications, conservative decision-making that allows for abstention can be advantageous. To this end, we introduce a novel approach that determines the optimal cutoff interval for risk scores, which can be…

Machine Learning · Statistics 2025-10-01 Yishu Wei , Wen-Yee Lee , George Ekow Quaye , Xiaogang Su

The interpretation of medical images is a challenging task, often complicated by the presence of artifacts, occlusions, limited contrast and more. Most notable is the case of chest radiography, where there is a high inter-rater variability…

In this paper, we investigate the hypothesis testing problem that checks whether part of covariates / confounders significantly affect the heterogeneous treatment effect given all covariates. This model checking is particularly useful in…

Statistics Theory · Mathematics 2020-09-24 Niwen Zhou , Xu Guo , Lixing Zhu

Estimating the individuals' potential response to varying treatment doses is crucial for decision-making in areas such as precision medicine and management science. Most recent studies predict counterfactual outcomes by learning a covariate…

Machine Learning · Computer Science 2024-03-22 Minqin Zhu , Anpeng Wu , Haoxuan Li , Ruoxuan Xiong , Bo Li , Xiaoqing Yang , Xuan Qin , Peng Zhen , Jiecheng Guo , Fei Wu , Kun Kuang

This paper illustrates the use of selected robust estimators of covariance or correlation in the identification of anomalous laboratory results in inter-laboratory data. It is shown that robust estimators can substantially reduce the impact…

Applications · Statistics 2019-05-29 Stephen L R Ellison