English
Related papers

Related papers: Detecting approximate replicate components of a hi…

200 papers

Sparse latent multi-factor models have been used in many exploratory and predictive problems with high-dimensional multivariate observations. Because of concerns with identifiability, the latent factors are almost always assumed to be…

Applications · Statistics 2013-12-09 Vinicius Diniz Mayrink , Joseph Edward Lucas

In this work we address the problem of approximating high-dimensional data with a low-dimensional representation. We make the following contributions. We propose an inverse regression method which exchanges the roles of input and response,…

Machine Learning · Computer Science 2015-09-04 Antoine Deleforge , Florence Forbes , Radu Horaud

We here introduce a novel classification approach adopted from the nonlinear model identification framework, which jointly addresses the feature selection and classifier design tasks. The classifier is constructed as a polynomial expansion…

Machine Learning · Computer Science 2016-07-29 Aida Brankovic , Alessandro Falsone , Maria Prandini , Luigi Piroddi

Single-Index Models are high-dimensional regression problems with planted structure, whereby labels depend on an unknown one-dimensional projection of the input via a generic, non-linear, and potentially non-deterministic transformation. As…

Machine Learning · Computer Science 2024-03-14 Alex Damian , Loucas Pillaud-Vivien , Jason D. Lee , Joan Bruna

Cognitive Diagnosis Models (CDMs) are a special family of discrete latent variable models that are widely used in modern educational, psychological, social and biological sciences. A key component of CDMs is a binary $Q$-matrix…

Methodology · Statistics 2025-01-08 Chenchen Ma , Gongjun Xu

Recovering low-rank structures via eigenvector perturbation analysis is a common problem in statistical machine learning, such as in factor analysis, community detection, ranking, matrix completion, among others. While a large variety of…

Statistics Theory · Mathematics 2019-05-06 Emmanuel Abbe , Jianqing Fan , Kaizheng Wang , Yiqiao Zhong

In this study, we develop a latent factor model for analysing high-dimensional binary data. Specifically, a standard probit model is used to describe the regression relationship between the observed binary data and the continuous latent…

Methodology · Statistics 2024-04-15 Jiaxin Shi , Yuan Gao , Rui Pan , Hansheng Wang

The spectra of random feature matrices provide essential information on the conditioning of the linear system used in random feature regression problems and are thus connected to the consistency and generalization of random feature models.…

Machine Learning · Statistics 2022-12-13 Zhijun Chen , Hayden Schaeffer , Rachel Ward

We study the estimation of a high dimensional approximate factor model in the presence of both cross sectional dependence and heteroskedasticity. The classical method of principal components analysis (PCA) does not efficiently estimate the…

Methodology · Statistics 2012-10-01 Jushan Bai , Yuan Liao

A novel framework is introduced to formalize identifiability in well-specified but ill-posed linear regression models. The framework is distribution-free and accommodates highly correlated features that may or may not relate to the…

Statistics Theory · Mathematics 2026-03-05 Gianluca Finocchio , Tatyana Krivobokova

Factor and sparse models are two widely used methods to impose a low-dimensional structure in high-dimensions. However, they are seemingly mutually exclusive. We propose a lifting method that combines the merits of these two models in a…

Econometrics · Economics 2022-09-07 Jianqing Fan , Ricardo Masini , Marcelo C. Medeiros

The latent class model is a widely used mixture model for multivariate discrete data. Besides the existence of qualitatively heterogeneous latent classes, real data often exhibit additional quantitative heterogeneity nested within each…

Methodology · Statistics 2025-01-23 Zhongyuan Lyu , Ling Chen , Yuqi Gu

A key challenge to performing effective analyses of high-dimensional data is finding a signal-rich, low-dimensional representation. For linear subspaces, this is generally performed by decomposing a design matrix (via eigenvalue or singular…

Computation · Statistics 2021-08-02 Wenlan Zang , Jen-hwa Chu , Michael J. Kane

We introduce and analyse a family of hash and predicate functions that are more likely to produce collisions for small reducible configurations of vectors. These may offer practical improvements to lattice sieving for short vectors. In…

Number Theory · Mathematics 2023-11-15 Gabriella Holden , Daniel Shiu , Lauren Strutt

In contemporary scientific research, it is of great interest to predict a categorical response based on a high-dimensional tensor (i.e. multi-dimensional array) and additional covariates. This mixture of different types of data leads to…

Methodology · Statistics 2018-05-14 Yuqing Pan , Qing Mai , Xin Zhang

We consider a discriminative learning (regression) problem, whereby the regression function is a convex combination of k linear classifiers. Existing approaches are based on the EM algorithm, or similar techniques, without provable…

Machine Learning · Computer Science 2014-08-01 Yuekai Sun , Stratis Ioannidis , Andrea Montanari

We present a greedy algorithm for computing selected eigenpairs of a large sparse matrix $H$ that can exploit localization features of the eigenvector. When the eigenvector to be computed is localized, meaning only a small number of its…

Computational Physics · Physics 2021-02-09 Taylor M. Hernandez , Roel Van Beeumen , Mark A. Caprio , Chao Yang

This work presents an adaptive group testing framework for the range-based high dimensional near neighbor search problem. Our method efficiently marks each item in a database as neighbor or non-neighbor of a query point, based on a cosine…

Data Structures and Algorithms · Computer Science 2024-09-10 Harsh Shah , Kashish Mittal , Ajit Rajwade

Parameterized systems of polynomial equations arise in many applications in science and engineering with the real solutions describing, for example, equilibria of a dynamical system, linkages satisfying design constraints, and scene…

Machine Learning · Statistics 2022-08-09 Edgar A. Bernal , Jonathan D. Hauenstein , Dhagash Mehta , Margaret H. Regan , Tingting Tang

Many tasks in data mining and related fields can be formalized as matching between objects in two heterogeneous domains, including collaborative filtering, link prediction, image tagging, and web search. Machine learning techniques,…

Machine Learning · Computer Science 2014-10-24 Jingbo Shang , Tianqi Chen , Hang Li , Zhengdong Lu , Yong Yu