中文
相关论文

相关论文: Deformed semicircle law and concentration of nonli…

200 篇论文

We consider the problem of learning an unknown function $f_{\star}$ on the $d$-dimensional sphere with respect to the square loss, given i.i.d. samples $\{(y_i,{\boldsymbol x}_i)\}_{i\le n}$ where ${\boldsymbol x}_i$ is a feature vector…

统计理论 · 数学 2020-02-18 Behrooz Ghorbani , Song Mei , Theodor Misiakiewicz , Andrea Montanari

This paper is concerned with the asymptotic distribution of the largest eigenvalues for some nonlinear random matrix ensemble stemming from the study of neural networks. More precisely we consider $M= \frac{1}{m} YY^\top$ with $Y=f(WX)$…

概率论 · 数学 2022-01-14 Lucas Benigni , Sandrine Péché

Many recent works have studied the eigenvalue spectrum of the Conjugate Kernel (CK) defined by the nonlinear feature map of a feedforward neural network. However, existing results only establish weak convergence of the empirical eigenvalue…

机器学习 · 统计学 2024-02-16 Zhichao Wang , Denny Wu , Zhou Fan

We consider inhomogeneous square random matrices of size $N$ with independent entries of mean 0 and finite variance. We assume that the variance profile of this matrix is doubly stochastic and has a band-like structure with an appropriately…

概率论 · 数学 2025-08-27 Yi Han

Modern deep networks are heavily overparameterized yet often generalize well, suggesting a form of low intrinsic complexity not reflected by parameter counts. We study this complexity at initialization through the effective rank of the…

机器学习 · 计算机科学 2025-12-02 Praveen Anilkumar Shukla

We study nonparametric regression by an over-parameterized two-layer neural network trained by gradient descent (GD) in this paper. We show that, if the neural network is trained by GD with early stopping, then the trained network renders a…

机器学习 · 统计学 2025-11-07 Yingzhen Yang , Ping Li

We study the optimization and sample complexity of gradient-based training of a two-layer neural network with quadratic activation function in the high-dimensional regime, where the data is generated as $f_*(\boldsymbol{x}) \propto…

机器学习 · 统计学 2026-01-01 Gérard Ben Arous , Murat A. Erdogdu , Nuri Mert Vural , Denny Wu

For two lacunary sequences $(M_{n,1})_{n\geq 2},(M_{n,2})_{n\geq 0}$ and suitable functions $f$ we introduce random matrix ensembles with \begin{equation*} X_{n,n'}=f(M_{n+n',1}x_1,M_{|n-n'|,2}x_2). \end{equation*} We prove weak convergence…

概率论 · 数学 2014-08-12 Thomas Löbbe

We compute the asymptotic eigenvalue distribution of the neural tangent kernel of a two-layer neural network under a specific scaling of dimension. Namely, if $X\in\mathbb{R}^{n\times d}$ is an i.i.d random matrix, $W\in\mathbb{R}^{d\times…

概率论 · 数学 2025-08-28 Lucas Benigni , Elliot Paquette

We study the eigenvalue distributions of the Conjugate Kernel and Neural Tangent Kernel associated to multi-layer feedforward neural networks. In an asymptotic regime where network width is increasing linearly in sample size, under random…

机器学习 · 统计学 2020-10-13 Zhou Fan , Zhichao Wang

This paper establishes rates of universal approximation for the shallow neural tangent kernel (NTK): network weights are only allowed microscopic changes from random initialization, which entails that activations are mostly unchanged, and…

机器学习 · 计算机科学 2020-02-18 Ziwei Ji , Matus Telgarsky , Ruicheng Xian

McKay proved that the limiting spectral measures of the ensembles of $d$-regular graphs with $N$ vertices converge to Kesten's measure as $N\to\infty$. In this paper we explore the case of weighted graphs. More precisely, given a large…

概率论 · 数学 2013-07-01 Leo Goldmakher , Cap Khoury , Steven J. Miller , Kesinee Ninsuwan

Double-descent curves in neural networks describe the phenomenon that the generalisation error initially descends with increasing parameters, then grows after reaching an optimal number of parameters which is less than the number of data…

机器学习 · 统计学 2023-05-29 Ouns El Harzli , Bernardo Cuenca Grau , Guillermo Valle-Pérez , Ard A. Louis

The paper is concerned with deformed Wigner random matrices. These matrices are closely related to Deep Neural Networks (DNNs): weight matrices of trained DNNs could be represented in the form $R + S$, where $R$ is random and $S$ is highly…

数学物理 · 物理学 2026-04-21 Ievgenii Afanasiev , Leonid Berlyand , Mariia Kiyashko

We address the structure identification and the uniform approximation of two fully nonlinear layer neural networks of the type $f(x)=1^T h(B^T g(A^T x))$ on $\mathbb R^d$ from a small number of query samples. We approach the problem by…

机器学习 · 计算机科学 2019-07-02 Massimo Fornasier , Timo Klock , Michael Rauchensteiner

We analyse the spectrum of additive finite-rank deformations of $N \times N$ Wigner matrices $H$. The spectrum of the deformed matrix undergoes a transition, associated with the creation or annihilation of an outlier, when an eigenvalue…

概率论 · 数学 2012-05-23 Antti Knowles , Jun Yin

Recent analyses of neural networks with shaped activations (i.e. the activation function is scaled as the network size grows) have led to scaling limits described by differential equations. However, these results do not a priori tell us…

机器学习 · 统计学 2024-04-22 Mufan Bill Li , Mihai Nica

Neural Tangent Kernel (NTK) theory is widely used to study the dynamics of infinitely-wide deep neural networks (DNNs) under gradient descent. But do the results for infinitely-wide networks give us hints about the behavior of real…

机器学习 · 计算机科学 2022-02-02 Mariia Seleznova , Gitta Kutyniok

Kernel ridge regression (KRR) is a popular class of machine learning models that has become an important tool for understanding deep learning. Much of the focus thus far has been on studying the proportional asymptotic regime, $n \asymp d$,…

机器学习 · 统计学 2025-10-07 Parthe Pandit , Zhichao Wang , Yizhe Zhu

In this paper, we study complex-valued neural network (CVNNs) with tensor-valued hidden-to-output weights within the framework of neural-network quantum field theory (NN-QFT). For standard CVNNs with scalar weights, we derive the generating…

高能物理 - 理论 · 物理学 2026-02-03 Guojun Huang , Kai Zhou
‹ 上一页 1 2 3 10 下一页 ›