English
Related papers

Related papers: LDLT L-Lipschitz Network Weight Parameterization I…

200 papers

Let $m$ be a bounded function and $\alpha$ a nonnegative parameter. This article is concerned with the first eigenvalue $\lambda\_\alpha(m)$ of the drifted Laplacian type operator $\mathcal L\_m$ given by $\mathcal L\_m(u)=…

Analysis of PDEs · Mathematics 2021-12-01 Idriss Mazari , Grégoire Nadin , Yannick Privat

Modern machine learning models are often trained in a setting where the number of parameters exceeds the number of training samples. To understand the implicit bias of gradient descent in such overparameterized models, prior work has…

Machine Learning · Statistics 2025-10-29 Hannes Matt , Dominik Stöger

Generative models that maximize model likelihood have gained traction in many practical settings. Among them, perturbation based approaches underpin many strong likelihood estimation models, yet they often face slow convergence and limited…

Information Theory · Computer Science 2025-10-27 Yirong Shen , Lu Gan , Cong Ling

This work proposes to reduce visibility data volume using a baseline-dependent lossy compression technique that preserves smearing at the edges of the field-of-view. We exploit the relation of the rank of a matrix and the fact that a…

Instrumentation and Methods for Astrophysics · Physics 2023-04-17 M Atemkeng , S Perkins , E Seck , S Makhathini , O Smirnov , L Bester , B Hugo

Overparameterized models may have many interpolating solutions; implicit regularization refers to the hidden preference of a particular optimization method towards a certain interpolating solution among the many. A by now established line…

Machine Learning · Computer Science 2024-09-18 Hung-Hsu Chou , Holger Rauhut , Rachel Ward

Recently mean field theory has been successfully used to analyze properties of wide, random neural networks. It gave rise to a prescriptive theory for initializing feed-forward neural networks with orthogonal weights, which ensures that…

Machine Learning · Statistics 2019-06-05 Piotr A. Sokol , Il Memming Park

Many network datasets exhibit connectivity with variance by resolution and large-scale organization that coexists with localized departures. When vertices have observed ordering or embedding, such as geography in spatial and village…

Statistics Theory · Mathematics 2025-12-23 Marios Papamichalis , Regina Ruane

A new initialization method for hidden parameters in a neural network is proposed. Derived from the integral representation of the neural network, a nonparametric probability distribution of hidden parameters is introduced. In this…

Machine Learning · Computer Science 2014-02-20 Sho Sonoda , Noboru Murata

It is challenging to develop stochastic gradient based scalable inference for deep discrete latent variable models (LVMs), due to the difficulties in not only computing the gradients, but also adapting the step sizes to different latent…

Machine Learning · Statistics 2017-06-07 Yulai Cong , Bo Chen , Hongwei Liu , Mingyuan Zhou

We investigate the sample complexity of bounded two-layer neural networks using different activation functions. In particular, we consider the class $$ \mathcal{H} = \left\{\textbf{x}\mapsto \langle \textbf{v}, \sigma \circ W\textbf{b} +…

Machine Learning · Computer Science 2024-01-23 Amit Daniely , Elad Granot

We propose a novel low-rank initialization framework for training low-rank deep neural networks -- networks where the weight parameters are re-parameterized by products of two low-rank matrices. The most successful prior existing approach,…

Machine Learning · Computer Science 2022-05-23 Kiran Vodrahalli , Rakesh Shivanna , Maheswaran Sathiamoorthy , Sagar Jain , Ed H. Chi

We consider the teacher-student setting of learning shallow neural networks with quadratic activations and planted weight matrix $W^*\in\mathbb{R}^{m\times d}$, where $m$ is the width of the hidden layer and $d\le m$ is the data dimension.…

Machine Learning · Statistics 2020-07-13 David Gamarnik , Eren C. Kızıldağ , Ilias Zadik

We study the multivariate nonparametric change point detection problem, where the data are a sequence of independent $p$-dimensional random vectors whose distributions are piecewise-constant with Lipschitz densities changing at unknown…

Statistics Theory · Mathematics 2020-06-26 Oscar Hernan Madrid Padilla , Yi Yu , Daren Wang , Alessandro Rinaldo

Deep residual networks (ResNets) have demonstrated outstanding success in computer vision tasks, attributed to their ability to maintain gradient flow through deep architectures. Simultaneously, controlling the Lipschitz bound in neural…

Machine Learning · Computer Science 2025-03-03 Marius F. R. Juston , William R. Norris , Dustin Nottage , Ahmet Soylemezoglu

Data augmentation is often used to incorporate inductive biases into models. Traditionally, these are hand-crafted and tuned with cross validation. The Bayesian paradigm for model selection provides a path towards end-to-end learning of…

Machine Learning · Statistics 2022-03-02 Pola Schwöbel , Martin Jørgensen , Sebastian W. Ober , Mark van der Wilk

Over-parameterized neural networks incur prohibitive memory and computational costs for resource-constrained deployment. The Strong Lottery Ticket (SLT) hypothesis suggests that randomly initialized networks contain sparse subnetworks…

Machine Learning · Computer Science 2026-03-11 Itamar Tsayag , Ofir Lindenbaum

As Einstein's equations for binary compact object inspiral have only been approximately or intermittently solved by analytic or numerical methods, the models used to infer parameters of gravitational wave (GW) sources are subject to…

General Relativity and Quantum Cosmology · Physics 2021-01-04 A. Z. Jan , A. B. Yelikar , J. Lange , R. O'Shaughnessy

We study the probability distribution function (PDF) of the smallest eigenvalue of Laguerre-Wishart matrices $W = X^\dagger X$ where $X$ is a random $M \times N$ ($M \geq N$) matrix, with complex Gaussian independent entries. We compute…

Mathematical Physics · Physics 2016-04-15 Anthony Perret , Gregory Schehr

We study Gaussian-copula models with discrete margins, with primary emphasis on low-count (Poisson) data. Our goal is exact yet computationally efficient maximum likelihood (ML) estimation in regimes where many observations contain small…

Methodology · Statistics 2025-11-18 Anna van Es , Eva Cantoni

We study reinforcement learning for episodic Markov Decision Processes (MDPs) whose transitions are modelled by a multinomial logistic (MNL) model. Existing algorithms for MNL mixture MDPs yield a regret of $\smash{\tilde{O}(dH^2\sqrt{T})}$…

Artificial Intelligence · Computer Science 2026-05-20 Pierre Boudart , Pierre Gaillard , Alessandro Rudi
‹ Prev 1 4 5 6 7 8 10 Next ›