English
Related papers

Related papers: Continuous Specialization Transition in the Soft C…

200 papers

This work focuses on the analysis of fully connected feed forward ReLU neural networks as they approximate a given, smooth function. In contrast to conventionally studied universal approximation properties under increasing architectures,…

Machine Learning · Computer Science 2024-06-24 Erion Morina , Martin Holler

The softmax gating function is arguably the most popular choice in mixture of experts modeling. Despite its widespread use in practice, the softmax gating may lead to unnecessary competition among experts, potentially causing the…

Machine Learning · Statistics 2024-11-05 Huy Nguyen , Nhat Ho , Alessandro Rinaldo

In spite of finite dimension ReLU neural networks being a consistent factor behind recent deep learning successes, a theory of feature learning in these models remains elusive. Currently, insightful theories still rely on assumptions…

Machine Learning · Computer Science 2025-04-01 Devon Jarvis , Richard Klein , Benjamin Rosman , Andrew M. Saxe

In a neural network with ReLU activations, the number of piecewise linear regions in the output can grow exponentially with depth. However, this is highly unlikely to happen when the initial parameters are sampled randomly, which therefore…

Machine Learning · Computer Science 2025-10-17 Max Milkert , David Hyde , Forrest Laine

We consider the optimization problem associated with fitting two-layer ReLU networks with $k$ hidden neurons, where labels are assumed to be generated by a (teacher) neural network. We leverage the rich symmetry exhibited by such models to…

Machine Learning · Computer Science 2021-09-22 Yossi Arjevani , Michael Field

Deep learning has been widely used in many fields, but the model training process usually consumes massive computational resources and time. Therefore, designing an efficient neural network training method with a provable convergence…

Machine Learning · Computer Science 2023-07-14 Lianke Qin , Zhao Song , Yuanyuan Yang

We study the renormalization of (softly) broken supersymmetric theories at the one loop level in detail. We perform this analysis in a superspace approach in which the supersymmetry breaking interactions are parameterized using spurion…

High Energy Physics - Theory · Physics 2008-11-26 Stefan Groot Nibbelink , Tino S. Nyawelo

We give the first dimension-efficient algorithms for learning Rectified Linear Units (ReLUs), which are functions of the form $\mathbf{x} \mapsto \max(0, \mathbf{w} \cdot \mathbf{x})$ with $\mathbf{w} \in \mathbb{S}^{n-1}$. Our algorithm…

Machine Learning · Computer Science 2016-12-06 Surbhi Goel , Varun Kanade , Adam Klivans , Justin Thaler

The permutation model is a classical spin system where elements of the symmetric group interact with one another. The partition function of this model is directly related to the entanglement structure of random quantum circuits and random…

Statistical Mechanics · Physics 2026-05-26 Ryuki Ito , Taisei Matsuo , Masayuki Ohzeki

We derive and solve flow equations for a general O(N)-symmetric effective potential including wavefunction renormalization corrections combined with a heat-kernel regularization. We investigate the model at finite temperature and study the…

High Energy Physics - Phenomenology · Physics 2009-10-31 O. Bohr , B. -J. Schaefer , J. Wambach

Despite their widespread success, deep neural networks remain critically vulnerable to adversarial attacks, posing significant risks in safety-sensitive applications. This paper investigates activation functions as a crucial yet…

Machine Learning · Computer Science 2025-07-31 Yunrui Yu , Kafeng Wang , Hang Su , Jun Zhu

An activation function has a significant impact on the efficiency and robustness of the neural networks. As an alternative, we evolved a cutting-edge non-monotonic activation function, Negative Stimulated Hybrid Activation Function (Nish).…

Machine Learning · Computer Science 2022-12-20 Yildiray Anagun , Sahin Isik

ReLU neural networks trained as surrogate models can be embedded exactly in mixed-integer linear programs (MILPs), enabling global optimization over the learned function. The tractability of the resulting MILP depends on structural…

Optimization and Control · Mathematics 2026-04-27 Calvin Tsay

Rectified Linear Units (ReLU) are the default choice for activation functions in deep neural networks. While they demonstrate excellent empirical performance, ReLU activations can fall victim to the dead neuron problem. In these cases, the…

Machine Learning · Computer Science 2023-02-14 Tim Whitaker , Darrell Whitley

Classical results in neural network approximation theory show how arbitrary continuous functions can be approximated by networks with a single hidden layer, under mild assumptions on the activation function. However, the classical theory…

Optimization and Control · Mathematics 2023-04-06 Tyler Lekang , Andrew Lamperski

We study the loss surface of a feed-forward neural network with ReLU non-linearities, regularized with weight decay. We show that the regularized loss function is piecewise strongly convex on an important open set which contains, under some…

Neural and Evolutionary Computing · Computer Science 2019-12-10 Tristan Milne

We continue the study of the renormalization group and decoupling of massive fields in curved space, started in the previous article and analyse the higher derivative sector of the vacuum metric-dependent action of the Standard Model. The…

High Energy Physics - Phenomenology · Physics 2009-11-10 Eduard V. Gorbar , Ilya L. Shapiro

Exponential Linear Units (ELUs) are a useful rectifier for constructing deep learning architectures, as they may speed up and otherwise improve learning by virtue of not have vanishing gradients and by having mean activations near zero.…

Machine Learning · Computer Science 2017-04-26 Jonathan T. Barron

With the continuous growth of neural network scales, low-precision quantization is widely used in edge accelerators. Classic multi-threshold activation hardware requires 2^n thresholds for n-bit outputs, causing a rapid increase in hardware…

Hardware Architecture · Computer Science 2026-02-27 Yuhao Liu , Salim Ullah , Akash Kumar

The successful training of neural networks hinges on the use of first order optimization methods, yet the theoretical characterization of these methods remains incomplete. This is especially true in settings with mild overparameterization.…

Machine Learning · Computer Science 2026-05-27 James Town , Etienne Boursier , Ben Lewis , Matthias Englert , Ranko Lazic