中文
相关论文

相关论文: The Neural Covariance SDE: Shaped Infinite Depth-a…

200 篇论文

For a large class of feature maps we provide a tight asymptotic characterisation of the test error associated with learning the readout layer, in the high-dimensional limit where the input dimension, hidden layer widths, and number of…

机器学习 · 统计学 2024-06-11 Dominik Schröder , Daniil Dmitriev , Hugo Cui , Bruno Loureiro

The approximation power of general feedforward neural networks with piecewise linear activation functions is investigated. First, lower bounds on the size of a network are established in terms of the approximation error and network depth…

机器学习 · 计算机科学 2018-07-02 Mohammad Mehrabi , Aslan Tchamkerten , Mansoor I. Yousefi

Deep linear networks have been extensively studied, as they provide simplified models of deep learning. However, little is known in the case of finite-width architectures with multiple outputs and convolutional layers. In this manuscript,…

The paper deals with distribution of singular values of product of random matrices arising in the analysis of deep neural networks. The matrices resemble the product analogs of the sample covariance matrices, however, an important…

数学物理 · 物理学 2020-11-23 Leonid Pastur

We study the geometric properties of random neural networks by investigating the boundary volumes of their excursion sets for different activation functions, as the depth increases. More specifically, we show that, for activations which are…

概率论 · 数学 2026-01-29 Simmaco Di Lillo , Domenico Marinucci , Michele Salvi , Stefano Vigogna

In this paper, we consider fully connected feed-forward deep neural networks where weights and biases are independent and identically distributed according to Gaussian distributions. Extending previous results (Matthews et al., 2018a;b;…

概率论 · 数学 2024-12-02 Daniele Bracale , Stefano Favaro , Sandra Fortini , Stefano Peluchetti

Deep neural networks (DNNs) at convergence consistently represent the training data in the last layer via a highly symmetric geometric structure referred to as neural collapse. This empirical evidence has spurred a line of theoretical…

机器学习 · 计算机科学 2024-10-08 Arthur Jacot , Peter Súkeník , Zihan Wang , Marco Mondelli

Achieving transparency in black-box deep learning algorithms is still an open challenge. High dimensional features and decisions given by deep neural networks (NN) require new algorithms and methods to expose its mechanisms. Current…

机器学习 · 计算机科学 2020-06-12 Schyler C. Sun , Chen Li , Zhuangkun Wei , Antonios Tsourdos , Weisi Guo

This paper aims to examine the characteristics of the posterior distribution of covariance/precision matrices in a "large $p$, large $n$" scenario, where $p$ represents the number of variables and $n$ is the sample size. Our analysis…

统计理论 · 数学 2026-02-02 Partha Sarkar , Kshitij Khare , Malay Ghosh , Matt P. Wand

This paper presents a novel technique based on gradient boosting to train the final layers of a neural network (NN). Gradient boosting is an additive expansion algorithm in which a series of models are trained sequentially to approximate a…

机器学习 · 计算机科学 2023-05-05 Seyedsaman Emami , Gonzalo Martínez-Muñoz

Smooth activation functions are ubiquitous in modern deep learning, yet their theoretical advantages over non-smooth counterparts remain poorly understood. In this work, we study both approximation and statistical properties of neural…

机器学习 · 统计学 2026-03-03 Yuhao Liu , Zilin Wang , Lei Wu , Shaobo Zhang

We present Shape-Tailored Deep Neural Networks (ST-DNN). ST-DNN extend convolutional networks (CNN), which aggregate data from fixed shape (square) neighborhoods, to compute descriptors defined on arbitrarily shaped regions. This is natural…

计算机视觉与模式识别 · 计算机科学 2021-02-18 Naeemullah Khan , Angira Sharma , Ganesh Sundaramoorthi , Philip H. S. Torr

Inference in deep Bayesian neural networks is only fully understood in the infinite-width limit, where the posterior flexibility afforded by increased depth washes out and the posterior predictive collapses to a shallow Gaussian process.…

机器学习 · 计算机科学 2022-05-03 Jacob A. Zavatone-Veth , Cengiz Pehlevan

We empirically demonstrate that full-batch gradient descent on neural network training objectives typically operates in a regime we call the Edge of Stability. In this regime, the maximum eigenvalue of the training loss Hessian hovers just…

机器学习 · 计算机科学 2022-11-24 Jeremy M. Cohen , Simran Kaur , Yuanzhi Li , J. Zico Kolter , Ameet Talwalkar

Stochastic epidemic models on networks are inherently high-dimensional and the resulting exact models are intractable numerically even for modest network sizes. Mean-field models provide an alternative but can only capture average…

种群与进化 · 定量生物学 2020-07-03 Francesco Di Lauro , Jean-Charles Croix , Luc Berthouze , István Kiss

We are interested in modeling networks in which the connectivity among the nodes and node attributes are random variables and interact with each other. We propose a probabilistic model that allows one to formulate jointly a probability…

概率论 · 数学 2016-09-07 Haiyan Cai

Spike-timing dependent plasticity (STDP) is an organizing principle of biological neural networks. While synchronous firing of neurons is considered to be an important functional block in the brain, how STDP shapes neural networks possibly…

神经元与认知 · 定量生物学 2009-05-20 Yuko K. Takahashi , Hiroshi Kori , Naoki Masuda

We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient descent can achieve optimality \emph{without} massive…

机器学习 · 计算机科学 2024-10-11 Semih Cayci , Atilla Eryilmaz

We study two-banded, non-Hermitian random matrices inspired by sparse neural networks with a circular, 1d topology. We focus on two paradigmatic models, an SSH chain and a ladder model, which have both non-Hermitian directional bias and…

无序系统与神经网络 · 物理学 2026-05-20 Richard Huang , David R. Nelson

Expanding neural networks during training is a promising way to augment capacity without retraining larger models from scratch. However, newly added neurons often fail to adjust to a trained network and become inactive, providing no…

机器学习 · 计算机科学 2025-09-24 Nikolas Chatzis , Ioannis Kordonis , Manos Theodosis , Petros Maragos