中文
相关论文

相关论文: MLPs at the EOC: Concentration of the NTK

200 篇论文

We propose a neural-network construction of Euclidean scalar quantum field theories from transformer attention heads, defining $n$-point correlators by averaging over random network parameters in the NN-QFT framework. For a single attention…

机器学习 · 计算机科学 2026-02-12 Dmitry S. Ageev , Yulia A. Ageeva

This study explores the influence of a Stark-like perturbative potential on a quantum particle confined to a cylindrical surface (QPCS) and its implications for extra-dimensional theories. The QPCS framework is particularly relevant to…

量子物理 · 物理学 2025-12-30 Deriyan Senjaya

Krylov complexity, as a novel measure of operator complexity under Heisenberg evolution, exhibits many interesting universal behaviors and also bounds many other complexity measures. In this work, we study Krylov complexity $\mathcal{K}(t)$…

高能物理 - 理论 · 物理学 2024-01-01 Haifeng Tang

Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models. In this paper, we study the evolution of tokens in deep encoder-only transformers at…

偏微分方程分析 · 数学 2026-05-12 Albert Alcalde , Leon Bungert , Konstantin Riedl , Tim Roith

Common infinite-width architectures such as Neural Tangent Kernels (NTKs) have historically shown weak performance compared to finite models. This is usually attributed to the absence of feature learning. We show that this explanation is…

机器学习 · 计算机科学 2024-10-25 Luke Sernau

Attention-based architectures have become ubiquitous in machine learning, yet our understanding of the reasons for their effectiveness remains limited. This work proposes a new way to understand self-attention networks: we show that their…

机器学习 · 计算机科学 2023-08-02 Yihe Dong , Jean-Baptiste Cordonnier , Andreas Loukas

The vanilla self-attention mechanism in Transformers can be viewed as a two-layer fast-weight MLP, whose weights are dynamically induced by inputs and whose hidden dimension is equal to the sequence length $N$. As the context extends, the…

机器学习 · 计算机科学 2026-05-12 Qishuai Wen , Zhiyuan Huang , Xianghan Meng , Wei He , Chun-Guang Li

Attention-based transformers have been remarkably successful at modeling generative processes across various domains and modalities. In this paper, we study the behavior of transformers on data drawn from \kth Markov processes, where the…

机器学习 · 计算机科学 2024-07-26 Nived Rajaraman , Marco Bondaschi , Kannan Ramchandran , Michael Gastpar , Ashok Vardhan Makkuva

We investigate the asymptotic behavior of the number of parts $K_n$ in the Ewens--Pitman partition model under the regime where the diversity parameter is scaled linearly with the sample size, that is, $\theta = \lambda n$ for some~$\lambda…

概率论 · 数学 2025-03-26 Rodrigo Ribeiro

In recent years, the neural tangent kernel (NTK) and neural network Gaussian process kernel (NNGP) have given theoreticians tractable limiting cases of fully connected neural networks. However, the property of these kernels are poorly…

机器学习 · 统计学 2026-04-28 David Holzmüller , Max Schölpple

Classical optimization theory establishes that zeroth-order (ZO) algorithms suffer from a dimension-dependent slowdown, with convergence rates typically scaling with the model dimension compared to first-order methods. However, in contrast…

机器学习 · 计算机科学 2026-05-06 Zhe Li , Bicheng Ying , Zidong Liu , Haibo Yang

We study decentralized multiagent optimization over networks, modeled as undirected graphs. The optimization problem consists of minimizing a nonconvex smooth function plus a convex extended-value function, which enforces constraints or…

最优化与控制 · 数学 2024-12-13 Xiaokai Chen , Tianyu Cao , Gesualdo Scutari

Recently, theoretical analyses of deep neural networks have broadly focused on two directions: 1) Providing insight into neural network training by SGD in the limit of infinite hidden-layer width and infinitesimally small learning rate…

机器学习 · 计算机科学 2023-09-27 Rajat Vadiraj Dwaraknath , Tolga Ergen , Mert Pilanci

We consider gradient-based optimisation of wide, shallow neural networks, where the output of each hidden node is scaled by a positive parameter. The scaling parameters are non-identical, differing from the classical Neural Tangent Kernel…

机器学习 · 统计学 2025-02-19 Francois Caron , Fadhel Ayed , Paul Jung , Hoil Lee , Juho Lee , Hongseok Yang

The multi-head attention layer is one of the key components of the transformer architecture that sets it apart from traditional feed-forward models. Given a sequence length $k$, attention matrices…

机器学习 · 计算机科学 2024-02-07 Sitan Chen , Yuanzhi Li

Despite their immense promise in performing a variety of learning tasks, a theoretical understanding of the limitations of Deep Neural Networks (DNNs) has so far eluded practitioners. This is partly due to the inability to determine the…

机器学习 · 计算机科学 2024-01-25 Saad Qadeer , Andrew Engel , Amanda Howard , Adam Tsou , Max Vargas , Panos Stinis , Tony Chiang

Are neural networks biased toward simple functions? Does depth always help learn more complex features? Is training the last layer of a network as good as training all layers? How to set the range for learning rate tuning? These questions…

机器学习 · 计算机科学 2020-04-10 Greg Yang , Hadi Salman

Zeroth-order (ZO) optimization enables memory-efficient training of neural networks by estimating gradients via forward passes only, eliminating the need for backpropagation. However, the stochastic nature of gradient estimation…

机器学习 · 计算机科学 2026-03-24 Chen Zhang , Yuxin Cheng , Chenchen Ding , Shuqi Wang , Jingreng Lei , Runsheng Yu , Yik-Chung WU , Ngai Wong

A new atmospheric neutrino oscillation tool which utilizes full three neutrino oscillation probabilities and a full three neutrino treatment of the MSW effect is combined with a standard analysis of the K2K, MINOS, and CHOOZ data to examine…

核理论 · 物理学 2009-08-20 J. E. Roa , D. C. Latimer , D. J. Ernst

Fidelity-based quantum kernels provide a direct interface between quantum feature maps and classical kernel methods, but they can exhibit exponential concentration: with increasing system size or circuit expressivity, the Gram matrix…

量子物理 · 物理学 2026-02-19 Claudia Zendejas-Morales , Debashis Saikia , Utkarsh Singh