中文
相关论文

相关论文: Biwhitening Reveals the Rank of a Count Matrix

200 篇论文

In this paper, we develop new statistical theory for probabilistic principal component analysis models in high dimensions. The focus is the estimation of the noise variance, which is an important and unresolved issue when the number of…

统计理论 · 数学 2014-06-23 Damien Passemier , Zhaoyuan Li , Jian-Feng Yao

We develop a Bayesian methodology aimed at simultaneously estimating low-rank and row-sparse matrices in a high-dimensional multiple-response linear regression model. We consider a carefully devised shrinkage prior on the matrix of…

统计方法学 · 统计学 2019-04-10 Antik Chakraborty , Anirban Bhattacharya , Bani K. Mallick

We consider the problem of estimating a rank-one matrix in Gaussian noise under a probabilistic model for the left and right factors of the matrix. The probabilistic model can impose constraints on the factors including sparsity and…

信息论 · 计算机科学 2015-09-16 Alyson K. Fletcher , Sundeep Rangan

Symmetry arises often when learning from high dimensional data. For example, data sets consisting of point clouds, graphs, and unordered sets appear routinely in contemporary applications, and exhibit rich underlying symmetries.…

最优化与控制 · 数学 2025-02-06 Mateo Díaz , Dmitriy Drusvyatskiy , Jack Kendrick , Rekha R. Thomas

The last decade has seen a revolution in the theory and application of machine learning and pattern recognition. Through these advancements, variable ranking has emerged as an active and growing research area and it is now beginning to be…

计算机视觉与模式识别 · 计算机科学 2017-06-20 Giorgio Roffo

We study the statistical decision process of detecting the signal from a `signal+noise' type matrix model with an additive Wigner noise. We propose a hypothesis test based on the linear spectral statistics of the data matrix, which does not…

统计理论 · 数学 2021-03-05 Ji Hyung Jung , Hye Won Chung , Ji Oon Lee

High-dimensional inference refers to problems of statistical estimation in which the ambient dimension of the data may be comparable to or possibly even larger than the sample size. We study an instance of high-dimensional inference in…

统计理论 · 数学 2009-12-31 Sahand Negahban , Martin J. Wainwright

We study the problem of estimating a low-rank positive semidefinite (PSD) matrix from a set of rank-one measurements using sensing vectors composed of i.i.d. standard Gaussian entries, which are possibly corrupted by arbitrary outliers.…

信息论 · 计算机科学 2016-12-21 Yuanxin Li , Yue Sun , Yuejie Chi

We study symmetric spiked matrix models with respect to a general class of noise distributions. Given a rank-1 deformation of a random noise matrix, whose entries are independently distributed with zero mean and unit variance, the goal is…

数据结构与算法 · 计算机科学 2022-02-22 Jingqiu Ding , Samuel B. Hopkins , David Steurer

We present a scalable Bayesian model for low-rank factorization of massive tensors with binary observations. The proposed model has the following key properties: (1) in contrast to the models based on the logistic or probit likelihood,…

机器学习 · 统计学 2015-08-19 Changwei Hu , Piyush Rai , Lawrence Carin

Consider the problem of estimating a low-rank matrix when its entries are perturbed by Gaussian noise. If the empirical distribution of the entries of the spikes is known, optimal estimators that exploit this knowledge can substantially…

统计理论 · 数学 2019-08-08 Andrea Montanari , Ramji Venkataramanan

The problem of structured matrix estimation has been studied mostly under strong noise dependence assumptions. This paper considers a general framework of noisy low-rank-plus-sparse matrix recovery, where the noise matrix may come from any…

机器学习 · 统计学 2025-04-07 Jinhang Chai , Jianqing Fan

We consider the problem of estimating a rank-1 signal corrupted by structured rotationally invariant noise, and address the following question: how well do inference algorithms perform when the noise statistics is unknown and hence Gaussian…

信息论 · 计算机科学 2022-05-23 Jean Barbier , TianQi Hou , Marco Mondelli , Manuel Sáenz

This work explores non-negative low-rank matrix factorization based on regularized Poisson models (PF or "Poisson factorization" for short) for recommender systems with implicit-feedback data. The properties of Poisson likelihood allow a…

机器学习 · 计算机科学 2022-02-25 David Cortes

How can we discern whether the covariance operator of a stochastic process is of reduced rank, and if so, what its precise rank is? And how can we do so at a given level of confidence? This question is central to a great deal of methods for…

统计方法学 · 统计学 2020-08-11 Anirvan Chakraborty , Victor M. Panaretos

Recommender system is a widely adopted technology in a diversified class of product lines. Modern day recommender system approaches include matrix factorization, learning to rank and deep learning paradigms, etc. Unlike many other…

信息检索 · 计算机科学 2023-06-13 Hao Wang

Principal component analysis (PCA) is a foundational tool in modern data analysis, and a crucial step in PCA is selecting the number of components to keep. However, classical selection methods (e.g., scree plots, parallel analysis, etc.)…

统计理论 · 数学 2026-05-28 David Hong , Yue Sheng , Edgar Dobriban

Boolean matrix factorization (BMF) is a popular and powerful technique for inferring knowledge from data. The mining result is the Boolean product of two matrices, approximating the input dataset. The Boolean product is a disjunction of…

机器学习 · 计算机科学 2019-07-02 Sibylle Hess , Nico Piatkowski , Katharina Morik

We propose an algorithm to impute and forecast a time series by transforming the observed time series into a matrix, utilizing matrix estimation to recover missing values and de-noise observed entries, and performing linear regression to…

机器学习 · 计算机科学 2019-04-29 Anish Agarwal , Muhammad Jehangir Amjad , Devavrat Shah , Dennis Shen

The comparison of benchmark error sets is an essential tool for the evaluation of theories in computational chemistry. The standard ranking of methods by their Mean Unsigned Error is unsatisfactory for several reasons linked to the…

统计方法学 · 统计学 2020-09-29 Pascal Pernot , Andreas Savin