Related papers: Activation Saturation and Floquet Spectrum Collaps…
The stunning empirical successes of neural networks currently lack rigorous theoretical explanation. What form would such an explanation take, in the face of existing complexity-theoretic lower bounds? A first step might be to show that…
We consider a system of $N$ neurons, each spiking randomly with rate depending on its membrane potential. When a neuron spikes, its potential is reset to $0$ and all other neurons receive an additional amount $h/N$ of potential, where $ h >…
We explore convergence of deep neural networks with the popular ReLU activation function, as the depth of the networks tends to infinity. To this end, we introduce the notion of activation domains and activation matrices of a ReLU network.…
We provide the first global optimization landscape analysis of $Neural\;Collapse$ -- an intriguing empirical phenomenon that arises in the last-layer classifiers and features of neural networks during the terminal phase of training. As…
Time-periodic (Floquet) drive is a powerful method to engineer quantum phases of matter, including fundamentally non-equilibrium states that are impossible in static Hamiltonian systems. One characteristic example is the anomalous Floquet…
Self-sustained neural activity in the absence of ongoing external input is a fundamental feature of nervous system dynamics, yet the conditions under which it can emerge in biophysically grounded network models remain incompletely…
Calculations for infinite nuclear matter with realistic nucleon-nucleon interactions suggest that the isoscalar effective mass of a nucleon at the saturation density, m*/m, equals 0.8 +/- 0.1. This result is at variance with empirical data…
We study approximation limits of single-hidden-layer neural networks with analytic activation functions under global coefficient constraints. Under uniform $\ell^1$ bounds, or more generally sub-exponential growth of the coefficients, we…
Deep learning requires several design choices, such as the nodes' activation functions and the widths, types, and arrangements of the layers. One consideration when making these choices is the vanishing-gradient problem, which is the…
Modern practice for training classification deepnets involves a Terminal Phase of Training (TPT), which begins at the epoch where training error first vanishes; During TPT, the training error stays effectively zero while training loss is…
Electron-phonon coupling is believed to be responsible for many spectral anomalies in the cuprate superconductors. In particular, the $B_{1g}$ buckling mode of the oxygen ion in the $CuO_{2}$ plane has been proposed to be responsible for…
We propose the Hyperbolic Tangent Exponential Linear Unit (TeLU), a neural network hidden activation function defined as TeLU(x)=xtanh(exp(x)). TeLU's design is grounded in the core principles of key activation functions, achieving strong…
We find approximate solutions to the renormalization group equation which governs the quantum evolution of the effective theory for the Color Glass Condensate. This is a functional Fokker-Planck equation which generates in particular the…
We examine the closedness of sets of realized neural networks of a fixed architecture in Sobolev spaces. For an exactly $m$-times differentiable activation function $\rho$, we construct a sequence of neural networks $(\Phi_n)_{n \in…
Neural mass models are ubiquitous in large scale brain modelling. At the node level they are written in terms of a set of ODEs with a nonlinearity that is typically a sigmoidal shape. Using structural data from brain atlases they may be…
This paper explores the topological signatures of ReLU neural network activation patterns. We consider feedforward neural networks with ReLU activation functions and analyze the polytope decomposition of the feature space induced by the…
Empirical works show that for ReLU neural networks (NNs) with small initialization, input weights of hidden neurons (the input weight of a hidden neuron consists of the weight from its input layer to the hidden neuron and its bias term)…
Doubly-stochastic matrices (DSM) are increasingly utilized in structure-preserving deep architectures -- such as Optimal Transport layers and Sinkhorn-based attention -- to enforce numerical stability and probabilistic interpretability. In…
We here study a network of synaptic relations mingling excitatory and inhibitory neuron nodes that displays oscillations quite similar to electroencephalogram (EEG) brain waves, and identify abrupt variations brought about by swift synaptic…
We analyze the expressivity of a universal deep neural network that can be organized as a series of nested qubit rotations, accomplished by adjustable data re-uploads. While the maximal expressive power increases with the depth of the…