English
Related papers

Related papers: MLPs at the EOC: Concentration of the NTK

200 papers

Natural gradients have been widely studied from both theoretical and empirical perspectives, and it is commonly believed that natural gradients have advantages over standard (Euclidean) gradients in capturing the intrinsic geometric…

Machine Learning · Computer Science 2025-09-30 Qinxun Bai , Steven Rosenberg , Wei Xu

We study best-policy identification for finite-horizon risk-sensitive reinforcement learning under the entropic risk measure. Recent work established a constant gap in the exponential horizon dependence between lower and upper bounds on the…

Machine Learning · Computer Science 2026-05-14 Amer Essakine , Claire Vernade

Let $W_n= \frac{1}{\sqrt n} M_n$ be a Wigner matrix whose entries have vanishing third moment, normalized so that the spectrum is concentrated in the interval $[-2,2]$. We prove a concentration bound for $N_I = N_I(W_n)$, the number of…

Probability · Mathematics 2013-08-13 Terence Tao , Van Vu

While several studies confirmed that machine-learned potentials (MLPs) can provide accurate free energies for determining phase stabilities, the abilities of MLPs for efficiently constructing a full phase diagram of multi-component systems…

Computational Physics · Physics 2022-08-26 Kyeongpung Lee , Yutack Park , Seungwu Han

Modern large language models (LLMs) excel at tasks that require storing and retrieving knowledge, such as factual recall and question answering. Transformers are central to this capability because they can encode information during training…

Machine Learning · Statistics 2026-03-18 Nuri Mert Vural , Alberto Bietti , Mahdi Soltanolkotabi , Denny Wu

The performance of the data-dependent neural tangent kernel (NTK; Jacot et al. (2018)) associated with a trained deep neural network (DNN) often matches or exceeds that of the full network. This implies that DNN training via gradient…

Machine Learning · Computer Science 2025-05-22 Johannes Schwab , Bryan Kelly , Semyon Malamud , Teng Andrea Xu

In maximally chaotic quantum systems, a class of out-of-time-order correlators (OTOCs) saturate the Maldacena-Shenker-Stanford (MSS) bound on chaos. Recently, it has been shown that the same OTOCs must also obey an infinite set of…

High Energy Physics - Theory · Physics 2022-02-16 Sandipan Kundu

We establish explicit dynamics for neural networks whose training objective has a regularising term that constrains the parameters to remain close to their initial value. This keeps the network in a lazy training regime, where the dynamics…

Machine Learning · Statistics 2023-12-21 Eugenio Clerico , Benjamin Guedj

Self-attention mechanism is the key of the Transformer but often criticized for its computation demands. Previous token pruning works motivate their methods from the view of computation redundancy but still need to load the full network and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Sihao Lin , Pumeng Lyu , Dongrui Liu , Tao Tang , Xiaodan Liang , Andy Song , Xiaojun Chang

We explore the equivalence between neural networks and kernel methods by deriving the first exact representation of any finite-size parametric classification model trained with gradient descent as a kernel machine. We compare our exact…

Machine Learning · Computer Science 2023-08-10 Brian Bell , Michael Geyer , David Glickenstein , Amanda Fernandez , Juston Moore

A recent trend in explainable AI research has focused on surrogate modeling, where neural networks are approximated as simpler ML algorithms such as kernel machines. A second trend has been to utilize kernel functions in various…

Machine Learning · Computer Science 2024-03-13 Andrew Engel , Zhichao Wang , Natalie S. Frank , Ioana Dumitriu , Sutanay Choudhury , Anand Sarwate , Tony Chiang

Transformer self-attention can be interpreted as a gradient flow on the unit sphere, in which tokens evolve under softmax interaction potentials and tend to form clusters. While prior work has established clustering behavior for single-head…

Machine Learning · Computer Science 2026-05-11 Ayan Pendharkar

An interesting approach to analyzing neural networks that has received renewed attention is to examine the equivalent kernel of the neural network. This is based on the fact that a fully connected feedforward network with one hidden layer,…

Machine Learning · Computer Science 2018-06-04 Russell Tsuchida , Farbod Roosta-Khorasani , Marcus Gallagher

We investigate a two-dimensional statistical model of N charged particles interacting via logarithmic repulsion in the presence of an oppositely charged compact region K whose charge density is determined by its equilibrium potential at an…

Classical Analysis and ODEs · Mathematics 2012-07-04 Christopher D. Sinclair , Maxim L. Yattselev

Variational wavefunctions offer a practical route around the exponential complexity of many-body Hilbert spaces, but their expressive power is often sharply constrained. Matrix product states, for instance, are efficient but limited to area…

Quantum Physics · Physics 2026-03-26 Nisarga Paul

Quantum kernel methods (QKMs) offer an appealing framework for machine learning on near-term quantum computers. However, QKMs generically suffer from exponential concentration, requiring an exponential number of measurements to resolve the…

Strongly Correlated Electrons · Physics 2025-08-15 Ayana Sarkar , Martin Schnee , Roya Radgohar , Mojde Fadaie , Victor Drouin-Touchette , Stefanos Kourtis

Understanding the fundamental principles behind the massive success of neural networks is one of the most important open questions in deep learning. However, due to the highly complex nature of the problem, progress has been relatively…

Machine Learning · Computer Science 2021-12-13 Lechao Xiao

We compute the asymptotic eigenvalue distribution of the neural tangent kernel of a two-layer neural network under a specific scaling of dimension. Namely, if $X\in\mathbb{R}^{n\times d}$ is an i.i.d random matrix, $W\in\mathbb{R}^{d\times…

Probability · Mathematics 2025-08-28 Lucas Benigni , Elliot Paquette

Self-attention is usually described as a flexible, content-adaptive way to mix a token with information from its past. We reinterpret causal self-attention transformers, the backbone of modern foundation models, within a probabilistic…

Machine Learning · Computer Science 2026-03-24 Deepak Agarwal , Dhyey Dharmendrakumar Mavani , Suyash Gupta , Karthik Sethuraman , Tejas Dharamsi

The goal of this paper is to study operators of the form, \[ Tf(x)= \psi(x)\int f(\gamma_t(x))K(t)\: dt, \] where $\gamma$ is a real analytic function defined on a neighborhood of the origin in $(t,x)\in \R^N\times \R^n$, satisfying…

Classical Analysis and ODEs · Mathematics 2011-05-24 Elias M. Stein , Brian Street
‹ Prev 1 8 9 10 Next ›