中文
相关论文

相关论文: MLPs at the EOC: Concentration of the NTK

200 篇论文

Natural gradients have been widely studied from both theoretical and empirical perspectives, and it is commonly believed that natural gradients have advantages over standard (Euclidean) gradients in capturing the intrinsic geometric…

机器学习 · 计算机科学 2025-09-30 Qinxun Bai , Steven Rosenberg , Wei Xu

We study best-policy identification for finite-horizon risk-sensitive reinforcement learning under the entropic risk measure. Recent work established a constant gap in the exponential horizon dependence between lower and upper bounds on the…

机器学习 · 计算机科学 2026-05-14 Amer Essakine , Claire Vernade

Let $W_n= \frac{1}{\sqrt n} M_n$ be a Wigner matrix whose entries have vanishing third moment, normalized so that the spectrum is concentrated in the interval $[-2,2]$. We prove a concentration bound for $N_I = N_I(W_n)$, the number of…

概率论 · 数学 2013-08-13 Terence Tao , Van Vu

While several studies confirmed that machine-learned potentials (MLPs) can provide accurate free energies for determining phase stabilities, the abilities of MLPs for efficiently constructing a full phase diagram of multi-component systems…

计算物理 · 物理学 2022-08-26 Kyeongpung Lee , Yutack Park , Seungwu Han

Modern large language models (LLMs) excel at tasks that require storing and retrieving knowledge, such as factual recall and question answering. Transformers are central to this capability because they can encode information during training…

机器学习 · 统计学 2026-03-18 Nuri Mert Vural , Alberto Bietti , Mahdi Soltanolkotabi , Denny Wu

The performance of the data-dependent neural tangent kernel (NTK; Jacot et al. (2018)) associated with a trained deep neural network (DNN) often matches or exceeds that of the full network. This implies that DNN training via gradient…

机器学习 · 计算机科学 2025-05-22 Johannes Schwab , Bryan Kelly , Semyon Malamud , Teng Andrea Xu

In maximally chaotic quantum systems, a class of out-of-time-order correlators (OTOCs) saturate the Maldacena-Shenker-Stanford (MSS) bound on chaos. Recently, it has been shown that the same OTOCs must also obey an infinite set of…

高能物理 - 理论 · 物理学 2022-02-16 Sandipan Kundu

We establish explicit dynamics for neural networks whose training objective has a regularising term that constrains the parameters to remain close to their initial value. This keeps the network in a lazy training regime, where the dynamics…

机器学习 · 统计学 2023-12-21 Eugenio Clerico , Benjamin Guedj

Self-attention mechanism is the key of the Transformer but often criticized for its computation demands. Previous token pruning works motivate their methods from the view of computation redundancy but still need to load the full network and…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Sihao Lin , Pumeng Lyu , Dongrui Liu , Tao Tang , Xiaodan Liang , Andy Song , Xiaojun Chang

We explore the equivalence between neural networks and kernel methods by deriving the first exact representation of any finite-size parametric classification model trained with gradient descent as a kernel machine. We compare our exact…

机器学习 · 计算机科学 2023-08-10 Brian Bell , Michael Geyer , David Glickenstein , Amanda Fernandez , Juston Moore

A recent trend in explainable AI research has focused on surrogate modeling, where neural networks are approximated as simpler ML algorithms such as kernel machines. A second trend has been to utilize kernel functions in various…

Transformer self-attention can be interpreted as a gradient flow on the unit sphere, in which tokens evolve under softmax interaction potentials and tend to form clusters. While prior work has established clustering behavior for single-head…

机器学习 · 计算机科学 2026-05-11 Ayan Pendharkar

An interesting approach to analyzing neural networks that has received renewed attention is to examine the equivalent kernel of the neural network. This is based on the fact that a fully connected feedforward network with one hidden layer,…

机器学习 · 计算机科学 2018-06-04 Russell Tsuchida , Farbod Roosta-Khorasani , Marcus Gallagher

We investigate a two-dimensional statistical model of N charged particles interacting via logarithmic repulsion in the presence of an oppositely charged compact region K whose charge density is determined by its equilibrium potential at an…

经典分析与常微分方程 · 数学 2012-07-04 Christopher D. Sinclair , Maxim L. Yattselev

Variational wavefunctions offer a practical route around the exponential complexity of many-body Hilbert spaces, but their expressive power is often sharply constrained. Matrix product states, for instance, are efficient but limited to area…

量子物理 · 物理学 2026-03-26 Nisarga Paul

Quantum kernel methods (QKMs) offer an appealing framework for machine learning on near-term quantum computers. However, QKMs generically suffer from exponential concentration, requiring an exponential number of measurements to resolve the…

Understanding the fundamental principles behind the massive success of neural networks is one of the most important open questions in deep learning. However, due to the highly complex nature of the problem, progress has been relatively…

机器学习 · 计算机科学 2021-12-13 Lechao Xiao

We compute the asymptotic eigenvalue distribution of the neural tangent kernel of a two-layer neural network under a specific scaling of dimension. Namely, if $X\in\mathbb{R}^{n\times d}$ is an i.i.d random matrix, $W\in\mathbb{R}^{d\times…

概率论 · 数学 2025-08-28 Lucas Benigni , Elliot Paquette

Self-attention is usually described as a flexible, content-adaptive way to mix a token with information from its past. We reinterpret causal self-attention transformers, the backbone of modern foundation models, within a probabilistic…

The goal of this paper is to study operators of the form, \[ Tf(x)= \psi(x)\int f(\gamma_t(x))K(t)\: dt, \] where $\gamma$ is a real analytic function defined on a neighborhood of the origin in $(t,x)\in \R^N\times \R^n$, satisfying…

经典分析与常微分方程 · 数学 2011-05-24 Elias M. Stein , Brian Street
‹ 上一页 1 8 9 10 下一页 ›