English
Related papers

Related papers: Sufficient dimension reduction with additional inf…

200 papers

A theory of sufficient dimension reduction (SDR) is developed from an optimizational perspective. In our formulation of the problem, instead of dealing with raw data, we assume that our ground truth includes a mapping ${\mathbf f}: {\mathbb…

Machine Learning · Computer Science 2018-08-21 Rustem Takhanov

In our work, we propose a novel formulation for supervised dimensionality reduction based on a nonlinear dependency criterion called Statistical Distance Correlation, Szekely et. al. (2007). We propose an objective which is free of…

Machine Learning · Computer Science 2016-01-05 Praneeth Vepakomma , Chetan Tonde , Ahmed Elgammal

Under two-phase designs, the outcome and several covariates and confounders are measured in the first phase, and a new predictor of interest, which may be costly to collect, can be measured on a subsample in the second phase, without…

Methodology · Statistics 2023-05-12 Jingjing Zou , Lori B. Daniels , Karen Messer , Daniel Rabinowitz

Satellite-based retrieval has become a popular PM2.5 monitoring method currently. To improve the retrieval performance, multiple variables are usually introduced as auxiliary variable in addition to aerosol optical depth (AOD). Different…

Atmospheric and Oceanic Physics · Physics 2019-07-09 Qianqian Yang , Qiangqiang Yuan , Linwei Yue , Huanfeng Shen , Liangpei Zhang

Two-phase sampling is commonly adopted for reducing cost and improving estimation efficiency. In many two-phase studies, the outcome and some cheap covariates are observed for a large sample in Phase I, and expensive covariates are obtained…

Methodology · Statistics 2025-10-14 Qingning Zhou , Kin Yau Wong

High-dimensional big data appears in many research fields such as image recognition, biology and collaborative filtering. Often, the exploration of such data by classic algorithms is encountered with difficulties due to `curse of…

Machine Learning · Computer Science 2016-07-13 Amit Bermanis , Aviv Rotbart , Moshe Salhov , Amir Averbuch

We study regression discontinuity designs in which many predetermined covariates, possibly much more than the number of observations, can be used to increase the precision of treatment effect estimates. We consider a two-step estimator…

Econometrics · Economics 2022-05-06 Alexander Kreiß , Christoph Rothe

We address a classical problem in statistics: adding two-way interaction terms to a regression model. As the covariate dimension increases quadratically, we develop an estimator that adapts well to this increase, while providing accurate…

Methodology · Statistics 2023-09-26 Mark A. van de Wiel , Matteo Amestoy , Jeroen Hoogland

Unsupervised dimension selection is an important problem that seeks to reduce dimensionality of data, while preserving the most useful characteristics. While dimensionality reduction is commonly utilized to construct low-dimensional…

Machine Learning · Statistics 2018-11-01 Jayaraman J. Thiagarajan , Rushil Anirudh , Rahul Sridhar , Peer-Timo Bremer

Predictive modeling from high-dimensional genomic data is often preceded by a dimension reduction step, such as principal components analysis (PCA). However, the application of PCA is not straightforward for multi-source data, wherein…

Quantitative Methods · Quantitative Biology 2017-07-19 Adam Kaplan , Eric F. Lock

A novel general framework is proposed in this paper for dimension reduction in regression to fill the gap between linear and fully nonlinear dimension reduction. The main idea is to transform first each of the raw predictors monotonically,…

Methodology · Statistics 2014-01-03 Tao Wang , Xu Guo , Peirong Xu , Lixing Zhu

In this paper, we consider multivariate response regression models with high dimensional predictor variables. One way to model the correlation among the response variables is through the low rank decomposition of the coefficient matrix,…

Methodology · Statistics 2015-08-06 Ruiyan Luo , Xin Qi

Dimension reduction is increasingly applied to high-dimensional biomedical data to improve its interpretability. When datasets are reduced to two dimensions, each observation is assigned an x and y coordinates and is represented as a point…

Machine Learning · Computer Science 2024-04-01 Daniel B. Hier , Tayo Obafemi-Ajayi , Gayla R. Olbricht , Devin M. Burns , Sasha Petrenko , Donald C. Wunsch

In this paper, we address the problem of conducting statistical inference in settings involving large-scale data that may be high-dimensional and contaminated by outliers. The high volume and dimensionality of the data require distributed…

Machine Learning · Statistics 2022-11-30 Emadaldin Mozafari-Majd , Visa Koivunen

Mutual information (MI) is a general measure of statistical dependence with widespread application across the sciences. However, estimating MI between multi-dimensional variables is challenging because the number of samples necessary to…

Quantitative Methods · Quantitative Biology 2025-03-06 Gokul Gowri , Xiao-Kang Lun , Allon M. Klein , Peng Yin

Data collection costs can vary widely across variables in data science tasks. Two-phase designs can be employed to save data collection costs. This paper considers the two-phase studies where inexpensive variables are collected for all…

Methodology · Statistics 2025-12-04 Ruoyu Wang , Qihua Wang , Wang Miao

Gradient-based dimension reduction decreases the cost of Bayesian inference and probabilistic modeling by identifying maximally informative (and informed) low-dimensional projections of the data and parameters, allowing high-dimensional…

Computation · Statistics 2025-06-02 Ricardo Baptista , Michael Brennan , Youssef Marzouk

This paper proposes a doubly robust two-stage semiparametric difference-in-difference estimator for estimating heterogeneous treatment effects with high-dimensional data. Our new estimator is robust to model miss-specifications and allows…

Econometrics · Economics 2020-09-08 Yang Ning , Sida Peng , Jing Tao

We propose dimension reduction methods for sparse, high-dimensional multivariate response regression models. Both the number of responses and that of the predictors may exceed the sample size. Sometimes viewed as complementary, predictor…

Statistics Theory · Mathematics 2013-02-14 Florentina Bunea , Yiyuan She , Marten H. Wegkamp

Over the past decades, statisticians and machine-learning researchers have developed literally thousands of new tools for the reduction of high-dimensional data in order to identify the variables most responsible for a particular trait.…

Machine Learning · Statistics 2012-05-31 Chamont Wang , Jana Gevertz , Chaur-Chin Chen , Leonardo Auslender