English
Related papers

Related papers: Randomized Learning of the Second-Moment Matrix of…

200 papers

Let $M$ be a compact and connected smooth manifold endowed with a smooth action of a finite group $\Gamma$, and let $f$ be a $\Gamma$-invariant Morse function on $M$. We prove that the space of $\Gamma$-invariant Riemannian metrics on $M$…

Differential Geometry · Mathematics 2017-12-01 Ignasi Mundet i Riera

Given $n$ i.i.d. observations, we study the problem of estimating the spectrum of weighted Laplace operators of the form $\Delta_f=\Delta + \alpha \nabla \log f\cdot \nabla$, where $f$ is a positive probability density on a known compact…

Statistics Theory · Mathematics 2025-12-01 Yann Chaubet , Vincent Divol

Robust covariance estimation is the following, well-studied problem in high dimensional statistics: given $N$ samples from a $d$-dimensional Gaussian $\mathcal{N}(\boldsymbol{0}, \Sigma)$, but where an $\varepsilon$-fraction of the samples…

Data Structures and Algorithms · Computer Science 2020-06-25 Jerry Li , Guanghao Ye

We consider the problem of minimizing the sum of two convex functions: one is the average of a large number of smooth component functions, and the other is a general convex function that admits a simple proximal mapping. We assume the whole…

Optimization and Control · Mathematics 2014-03-20 Lin Xiao , Tong Zhang

We consider semidefinite programs (SDPs) of size n with equality constraints. In order to overcome scalability issues, Burer and Monteiro proposed a factorized approach based on optimizing over a matrix Y of size $n$ by $k$ such that $X =…

Machine Learning · Statistics 2018-11-29 Thomas Pumir , Samy Jelassi , Nicolas Boumal

The softmax activation function plays a crucial role in the success of large language models (LLMs), particularly in the self-attention mechanism of the widely adopted Transformer architecture. However, the underlying learning dynamics that…

Machine Learning · Computer Science 2026-01-27 Yang Cao , Yingyu Liang , Zhenmei Shi , Zhao Song

Gaussian smoothing (GS) is a derivative-free optimization (DFO) algorithm that estimates the gradient of an objective using perturbations of the current parameters sampled from a standard normal distribution. We generalize it to sampling…

Machine Learning · Computer Science 2022-11-29 Katelyn Gao , Ozan Sener

The stochastic mirror descent (SMD) algorithm is a general class of training algorithms, which includes the celebrated stochastic gradient descent (SGD), as a special case. It utilizes a mirror potential to influence the implicit bias of…

Machine Learning · Computer Science 2022-10-28 Taylan Kargin , Fariborz Salehi , Babak Hassibi

We study the optimization of non-convex functions that are not necessarily smooth (gradient and/or Hessian are Lipschitz) using first order methods. Smoothness is a restrictive assumption in machine learning in both theory and practice,…

Optimization and Control · Mathematics 2025-06-27 Daniel Yiming Cao , August Y. Chen , Karthik Sridharan , Benjamin Tang

A function $f: \mathbb{R}^d \rightarrow \mathbb{R}$ is a Sparse Additive Model (SPAM), if it is of the form $f(\mathbf{x}) = \sum_{l \in \mathcal{S}}\phi_{l}(x_l)$ where $\mathcal{S} \subset [d]$, $|\mathcal{S}| \ll d$. Assuming $\phi$'s,…

Machine Learning · Computer Science 2017-05-09 Hemant Tyagi , Anastasios Kyrillidis , Bernd Gärtner , Andreas Krause

In this paper we consider the unconstrained minimization problem of a smooth function in ${\mathbb{R}}^n$ in a setting where only function evaluations are possible. We design a novel randomized derivative-free algorithm --- the stochastic…

Optimization and Control · Mathematics 2019-05-08 El Houcine Bergou , Eduard Gorbunov , Peter Richtárik

We study the first gradient descent step on the first-layer parameters $\boldsymbol{W}$ in a two-layer neural network: $f(\boldsymbol{x}) = \frac{1}{\sqrt{N}}\boldsymbol{a}^\top\sigma(\boldsymbol{W}^\top\boldsymbol{x})$, where…

Machine Learning · Statistics 2022-05-04 Jimmy Ba , Murat A. Erdogdu , Taiji Suzuki , Zhichao Wang , Denny Wu , Greg Yang

We present an algorithm for approximating a function defined over a $d$-dimensional manifold utilizing only noisy function values at locations sampled from the manifold with noise. To produce the approximation we do not require any…

Machine Learning · Statistics 2020-08-13 Barak Sober , Yariv Aizenbud , David Levin

Let $B$ be a M\"obius band and $f:B \to \mathbb{R}$ be a Morse map taking a constant value on $\partial B$, and $\mathcal{S}(f,\partial B)$ be the group of diffeomorphisms $h$ of $B$ fixed on $\partial B$ and preserving $f$ in the sense…

Geometric Topology · Mathematics 2019-01-14 Iryna Kuznietsova , Sergiy Maksymenko

In this paper, we focus on the approximation of smooth functions $f: [-\pi, \pi] \rightarrow \mathbb{C}$, up to an unresolvable global phase ambiguity, from a finite set of Short Time Fourier Transform (STFT) magnitude (i.e., spectrogram)…

Numerical Analysis · Mathematics 2021-06-07 Mark Iwen , Michael Perlmutter , Nada Sissouno , Aditya Viswanathan

We consider network topology identification subject to a signal smoothness prior on the nodal observations. A fast dual-based proximal gradient algorithm is developed to efficiently tackle a strongly convex, smoothness-regularized network…

Machine Learning · Computer Science 2021-10-20 Seyed Saman Saboksayr , Gonzalo Mateos

In this paper, we consider a class of structured nonsmooth optimization problems over an embedded submanifold of a Euclidean space, where the first part of the objective is the sum of a difference-of-convex (DC) function and a smooth…

Optimization and Control · Mathematics 2025-11-07 Qia Li , Na Zhang , Junyu Feng , Hanwei Yan

$k$-subset sampling is ubiquitous in machine learning, enabling regularization and interpretability through sparsity. The challenge lies in rendering $k$-subset sampling amenable to end-to-end learning. This has typically involved relaxing…

Machine Learning · Computer Science 2024-06-10 Kareem Ahmed , Zhe Zeng , Mathias Niepert , Guy Van den Broeck

Information retrieval (IR) systems traditionally aim to maximize metrics built on rankings, such as precision or NDCG. However, the non-differentiability of the ranking operation prevents direct optimization of such metrics in…

Information Retrieval · Computer Science 2021-05-04 Thibaut Thonet , Yagmur Gizem Cinar , Eric Gaussier , Minghan Li , Jean-Michel Renders

We study the problem of approximating an unknown function $f:\mathbb{R}\to\mathbb{R}$ by a degree-$d$ polynomial using as few function evaluations as possible, where error is measured with respect to a probability distribution $\mu$.…

Data Structures and Algorithms · Computer Science 2025-08-11 Chris Camaño , Raphael A. Meyer , Kevin Shu
‹ Prev 1 8 9 10 Next ›