English
Related papers

Related papers: Optimal Subspace Estimation Using Overidentifying …

200 papers

We developed a statistical inference method applicable to a broad range of generalized linear models (GLMs) in high-dimensional settings, where the number of unknown coefficients scales proportionally with the sample size. Although a…

Statistics Theory · Mathematics 2024-05-24 Kazuma Sawaya , Yoshimasa Uematsu , Masaaki Imaizumi

We consider the problem of surrogate sufficient dimension reduction, that is, estimating the central subspace of a regression model, when the covariates are contaminated by measurement error. When no measurement error is present, a…

Methodology · Statistics 2023-10-24 Linh H. Nghiem , Francis K. C. Hui , Samuel Mueller , A. H. Welsh

We present Submatrix-wise Vector Embedding Learner (Swivel), a method for generating low-dimensional feature embeddings from a feature co-occurrence matrix. Swivel performs approximate factorization of the point-wise mutual information…

Computation and Language · Computer Science 2016-02-09 Noam Shazeer , Ryan Doherty , Colin Evans , Chris Waterson

Factor analysis provides a canonical framework for imposing lower-dimensional structure such as sparse covariance in high-dimensional data. High-dimensional data on the same set of variables are often collected under different conditions,…

Methodology · Statistics 2024-08-27 Noirrit Kiran Chandra , David B. Dunson , Jason Xu

Regression is one of the most fundamental statistical inference problems. A broad definition of regression problems is as estimation of the distribution of an outcome using a family of probability models indexed by covariates. Despite the…

Statistics Theory · Mathematics 2023-09-26 Peter Mueller , Fernando Andrés Quintana , Garritt L. Page

We consider low-distortion embeddings for subspaces under \emph{entrywise nonlinear transformations}. In particular we seek embeddings that preserve the norm of all vectors in a space $S = \{y: y = f(x)\text{ for }x \in Z\}$, where $Z$ is a…

Machine Learning · Computer Science 2020-10-09 Aarshvi Gajjar , Cameron Musco

Extracting latent low-dimensional structure from high-dimensional data is of paramount importance in timely inference tasks encountered with `Big Data' analytics. However, increasingly noisy, heterogeneous, and incomplete datasets as well…

Machine Learning · Statistics 2015-06-19 Morteza Mardani , Gonzalo Mateos , Georgios B. Giannakis

We describe a numerical scheme for evaluating the posterior moments of Bayesian linear regression models with partial pooling of the coefficients. The principal analytical tool of the evaluation is a change of basis from coefficient space…

Computation · Statistics 2021-10-01 Philip Greengard , Andrew Gelman , Aki Vehtari

The central subspace of a pair of random variables $(y,x) \in \mathbb{R}^{p+1}$ is the minimal subspace $\mathcal{S}$ such that $y \perp \hspace{-2mm} \perp x\mid P_{\mathcal{S}}x$. In this paper, we consider the minimax rate of estimating…

Statistics Theory · Mathematics 2017-01-25 Qian Lin , Xinran Li , Dongming Huang , Jun S. Liu

Small area models are mixed effects regression models that link the small areas and borrow strength from similar domains. When the auxiliary variables used in the models are measured with error, small area estimators that ignore the…

Methodology · Statistics 2018-10-23 Serena Arima , Silvia Polettini

We study optimization problems whereby the optimization variable is a probability measure. Since the probability space is not a vector space, many classical and powerful methods for optimization (e.g., gradients) are of little help. Thus,…

Optimization and Control · Mathematics 2024-06-18 Nicolas Lanzetti , Antonio Terpin , Florian Dörfler

We introduce an estimation method of covariance matrices in a high-dimensional setting, i.e., when the dimension of the matrix, , is larger than the sample size . Specifically, we propose an orthogonally equivariant estimator. The…

Statistics Theory · Mathematics 2020-12-04 Samprit Banerjee , Stefano Monni

Sliced inverse regression is a popular tool for sufficient dimension reduction, which replaces covariates with a minimal set of their linear combinations without loss of information on the conditional distribution of the response given the…

Machine Learning · Statistics 2018-09-18 Kean Ming Tan , Zhaoran Wang , Tong Zhang , Han Liu , R. Dennis Cook

Evaluating the statistical dimension is a common tool to determine the asymptotic phase transition in compressed sensing problems with Gaussian ensemble. Unfortunately, the exact evaluation of the statistical dimension is very difficult and…

Information Theory · Computer Science 2019-06-06 Sajad Daei , Farzan Haddadi , Arash Amini , Martin Lotz

The aim of this paper is to study the full $K-$moment problem for measures supported on some particular non-linear subsets $K$ of an infinite dimensional vector space. We focus on the case of random measures, that is $K$ is a subset of all…

Functional Analysis · Mathematics 2021-08-16 Maria Infusino , Tobias Kuna

In this paper, we propose a novel method for transforming data into a low-dimensional space optimized for one-class classification. The proposed method iteratively transforms data into a new subspace optimized for ellipsoidal encapsulation…

Machine Learning · Computer Science 2020-09-15 Fahad Sohrab , Jenni Raitoharju , Alexandros Iosifidis , Moncef Gabbouj

The state-of-the-art methods for estimating high-dimensional covariance matrices all shrink the eigenvalues of the sample covariance matrix towards a data-insensitive shrinkage target. The underlying shrinkage transformation is either…

Machine Learning · Statistics 2025-11-25 Man-Chung Yue , Yves Rychener , Daniel Kuhn , Viet Anh Nguyen

We consider quantile estimation in a semi-supervised setting, characterized by two available data sets: (i) a small or moderate sized labeled data set containing observations for a response and a set of possibly high dimensional covariates,…

Methodology · Statistics 2024-08-15 Abhishek Chakrabortty , Guorong Dai , Raymond J. Carroll

Data visualisation helps understanding data represented by multiple variables, also called features, stored in a large matrix where individuals are stored in lines and variable values in columns. These data structures are frequently called…

Human-Computer Interaction · Computer Science 2022-07-25 Haseeb Younis , Paul Trust , Rosane Minghim

Finding a small set of representatives from an unlabeled dataset is a core problem in a broad range of applications such as dataset summarization and information extraction. Classical exemplar selection methods such as $k$-medoids work…

Machine Learning · Computer Science 2020-06-09 Chong You , Chi Li , Daniel P. Robinson , Rene Vidal