Related papers: On subdifferential chain rule of matrix factorizat…
Despite its wide range of applications across various domains, the optimization foundations of deep matrix factorization (DMF) remain largely open. In this work, we aim to fill this gap by conducting a comprehensive study of the loss…
The study of a machine learning problem is in many ways is difficult to separate from the study of the loss function being used. One avenue of inquiry has been to look at these loss functions in terms of their properties as scoring rules…
Low rank matrix factorization is a fundamental building block in machine learning, used for instance to summarize gene expression profile data or word-document counts. To be robust to outliers and differences in scale across features, a…
We show, using detailed numerical analysis and theoretical arguments, that the normalized participation number of the stationary solutions of disordered nonlinear lattices obeys a one-parameter scaling law. Our approach opens a new way to…
The fluctuation-dissipation theorem (FDT) is a simple yet powerful consequence of the first-order differential equation governing the dynamics of systems subject simultaneously to dissipative and stochastic forces. The linear learning…
We consider deep linear networks with arbitrary convex differentiable loss. We provide a short and elementary proof of the fact that all local minima are global minima if the hidden layers are either 1) at least as wide as the input layer,…
We propose statistically robust and computationally efficient linear learning methods in the high-dimensional batch setting, where the number of features $d$ may exceed the sample size $n$. We employ, in a generic learning setting, two…
Matrix factorization techniques have been widely used as a method for collaborative filtering for recommender systems. In recent times, different variants of deep learning algorithms have been explored in this setting to improve the task of…
We study the problem of learning latent variables in Gaussian graphical models. Existing methods for this problem assume that the precision matrix of the observed variables is the superposition of a sparse and a low-rank component. In this…
We study the factor model problem, which aims to uncover low-dimensional structures in high-dimensional datasets. Adopting a robust data-driven approach, we formulate the problem as a saddle-point optimization. Our primary contribution is a…
The main purpose of this paper is the study of a~new class of summing multilinear operators acting from the product of Banach lattices with some nontrivial lattice convexity. A~mixed Pietsch-Maurey-Rosenthal type factorization theorem for…
For a proper extended real-valued function, this work focuses on the relationship between the subregularity of its subdifferential mapping relative to the critical set and its KL property of exponent 1/2. When the function is lsc convex, we…
In this short note, we derive an upper estimate of Clarke's subdifferential of marginal functions in Banach spaces. The structure of the upper estimate is very similar to other results already obtained in the literature. The novelty lies on…
In [2] we characterized in terms of a quadratic growth condition various metric regularity properties of the subdifferential of a lower semicontinuous convex function acting in a Hilbert space. Motivated by some recent results in [16] where…
We consider high order approximations of the solution of the stochastic filtering problem, derive their pathwise representation in the spirit of the earlier work of Clark and Davis and prove their robustness property. In particular, we show…
Lipschitz constants are connected to many properties of neural networks, such as robustness, fairness, and generalization. Existing methods for computing Lipschitz constants either produce relatively loose upper bounds or are limited to…
This paper considers a restriction to non-negative matrix factorization in which at least one matrix factor is stochastic. That is, the elements of the matrix factors are non-negative and the columns of one matrix factor sum to 1. This…
Designing learning algorithms that are resistant to perturbations of the underlying data distribution is a problem of wide practical and theoretical importance. We present a general approach to this problem focusing on unsupervised…
Given an element $f$ in a regular local ring, we study matrix factorizations of $f$ with $d \ge 2$ factors, that is, we study tuples of square matrices $(\varphi_1,\varphi_2,\dots,\varphi_d)$ such that their product is $f$ times an identity…
Investigating the dynamics of learning in machine learning algorithms is of paramount importance for understanding how and why an approach may be successful. The tools of physics and statistics provide a robust setting for such…