English
Related papers

Related papers: Determining the Dimension and Structure of the Sub…

200 papers

We consider the hypothesis testing problem of deciding whether an observed high-dimensional vector has independent normal components or, alternatively, if it has a small subset of correlated components. The correlated components may have a…

Statistics Theory · Mathematics 2012-06-04 Ery Arias-Castro , Sébastien Bubeck , Gábor Lugosi

We introduce a method to learn a hierarchy of successively more abstract representations of complex data based on optimizing an information-theoretic objective. Intuitively, the optimization searches for a set of latent factors that best…

Machine Learning · Computer Science 2014-11-03 Greg Ver Steeg , Aram Galstyan

Dimension reduction for high-dimensional compositional data plays an important role in many fields, where the principal component analysis of the basis covariance matrix is of scientific interest. In practice, however, the basis variables…

Methodology · Statistics 2021-09-13 Jingru Zhang , Wei Lin

In many scientific tasks we are interested in discovering whether there exist any correlations in our data. This raises many questions, such as how to reliably and interpretably measure correlation between a multivariate set of attributes,…

Machine Learning · Computer Science 2019-09-02 Panagiotis Mandros , Mario Boley , Jilles Vreeken

We propose a method for testing whether hierarchically ordered groups of potentially correlated variables are significant for explaining a response in a high-dimensional linear model. In presence of highly correlated variables, as is very…

Statistics Theory · Mathematics 2014-09-04 Jacopo Mandozzi , Peter Bühlmann

How does one find dimensions in multivariate data that are reliably expressed across repetitions? For example, in a brain imaging study one may want to identify combinations of neural signals that are reliably expressed across multiple…

Machine Learning · Statistics 2022-12-05 Lucas C. Parra , Stefan Haufe , Jacek P. Dmochowski

We study the problem of linear feature selection when features are highly correlated. Such settings pose two fundamental challenges. First, how should model similarity be defined? Simply counting features in common can be misleading: two…

Methodology · Statistics 2026-03-24 Xiaozhu Zhang , Jacob Bien , Armeen Taeb

We obtain general, exact formulas for the overlaps between the eigenvectors of large correlated random matrices, with additive or multiplicative noise. These results have potential applications in many different contexts, from quantum…

Statistical Mechanics · Physics 2018-12-05 Joël Bun , Jean-Philippe Bouchaud , Marc Potters

Statistical inference of the dependence between objects often relies on covariance matrices. Unless the number of features (e.g. data points) is much larger than the number of objects, covariance matrix cleaning is necessary to reduce…

Risk Management · Quantitative Finance 2021-06-09 Christian Bongiorno , Damien Challet

Multi-modal data is becoming more common in big data background. Finding the semantically similar objects from different modality is one of the heart problems of multi-modal learning. Most of the current methods try to learn the inter-modal…

Artificial Intelligence · Computer Science 2018-09-05 Qibin Zheng , Xingchun Diao , Jianjun Cao , Xiaolei Zhou , Yi Liu , Hongmei Li

We consider the problem of estimating the principal components of a population correlation matrix from a limited number of measurement data. Using a combination of random matrix and information-theoretic tools, we show that all the…

Statistical Mechanics · Physics 2016-01-20 Rémi Monasson , Dario Villamaina

Covariance matrices of random vectors contain information that is crucial for modelling. Specific structures and patterns of the covariances (or correlations) may be used to justify parametric models, e.g., autoregressive models. Until now,…

Methodology · Statistics 2025-02-11 Paavo Sattler , Dennis Dobler

Linear regression models depend directly on the design matrix and its properties. Techniques that efficiently estimate model coefficients by partitioning rows of the design matrix are increasingly popular for large-scale problems because…

Machine Learning · Statistics 2019-07-23 Michael J. Kane , Bryan Lewis , Sekhar Tatikonda , Simon Urbanek

Correlation networks derived from multivariate data appear in many applications across the sciences. These networks are usually dense and require sparsification to detect meaningful structure. However, current methods for sparsifying…

Physics and Society · Physics 2023-03-06 Magnus Neuman , Viktor Jonsson , Joaquín Calatayud , Martin Rosvall

Many complex tasks can be decomposed into simpler, independent parts. Discovering such underlying compositional structure has the potential to enable compositional generalization. Despite progress, our most powerful systems struggle to…

The development of science has been transforming man's view towards nature for centuries. Observing structures and patterns in an effective approach to discover regularities from data is a key step toward theory-building. With increasingly…

Computational Physics · Physics 2025-06-09 Guang-Xing Li

We investigate the problem of detecting dependencies between the components of a high-dimensional vector. Our approach advances the existing literature in two important respects. First, we consider the problem under privacy constraints.…

Statistics Theory · Mathematics 2026-03-24 Patrick Bastian , Holger Dette , Martin Dunsche

The assumption of independent subvectors arises in many aspects of multivariate analysis. In most real-world applications, however, we lack prior knowledge about the number of subvectors and the specific variables within each subvector.…

Methodology · Statistics 2024-01-23 Jan O. Bauer

Hypergraphs, increasingly utilised to model complex and diverse relationships in modern networks, have gained significant attention for representing intricate higher-order interactions. Among various challenges, cohesive subgraph discovery…

Social and Information Networks · Computer Science 2025-07-14 Dahee Kim , Hyewon Kim , Song Kim , Minseok Kim , Junghoon Kim , Yeon-Chang Lee , Sungsu Lim

Much of scientific data is collected as randomized experiments intervening on some and observing other variables of interest. Quite often, a given phenomenon is investigated in several studies, and different sets of variables are involved…

Methodology · Statistics 2012-10-19 Antti Hyttinen , Frederick Eberhardt , Patrik O. Hoyer
‹ Prev 1 2 3 10 Next ›