English
Related papers

Related papers: Can an MLP Absorb Its Own Skip Connection?

200 papers

We consider the problem of learning an arbitrarily-biased ReLU activation (or neuron) over Gaussian marginals with the squared loss objective. Despite the ReLU neuron being the basic building block of modern neural networks, we still do not…

Machine Learning · Computer Science 2026-02-04 Anxin Guo , Aravindan Vijayaraghavan

When can the input of a ReLU neural network be inferred from its output? In other words, when is the network injective? We consider a single layer, $x \mapsto \mathrm{ReLU}(Wx)$, with a random Gaussian $m \times n$ matrix $W$, in a…

Disordered Systems and Neural Networks · Physics 2024-12-13 Antoine Maillard , Afonso S. Bandeira , David Belius , Ivan Dokmanić , Shuta Nakajima

Length generalization remains a persistent challenge for neural networks: recurrent models tend to suffer from positional biases, while transformers are constrained by fixed computational depth. Regular languages provide a frequently used…

Machine Learning · Computer Science 2026-05-26 Charles Pert , Dalal Alrajeh , Alessandra Russo

Recent results in nonparametric regression show that deep learning, i.e., neural network estimates with many hidden layers, are able to circumvent the so-called curse of dimensionality in case that suitable restrictions on the structure of…

Machine Learning · Statistics 2020-09-30 Michael Kohler , Sophie Langer

Can a transformer learn which attention entries matter during training? In principle, yes: attention distributions are highly concentrated, and a small gate network can identify the important entries post-hoc with near-perfect accuracy. In…

Machine Learning · Computer Science 2026-03-04 Keston Aquino-Michaels

While Attention Residuals has shown some effectiveness in addressing the widespread issue of unbounded activation growth across deep residual layers, it inevitably incurs significant communication overhead. To circumvent this bottleneck, we…

Machine Learning · Computer Science 2026-05-25 Zhizhan Zheng , Feiyun Zhang , Shuchun Liu , Tian Xia , Xi Liu , Dasheng Hu , Hongquan Zhou

We analyze a simple one-hidden-layer neural network with ReLU activation functions and fixed biases, with one-dimensional input and output. We study both continuous and discrete versions of the model, and we rigorously prove the convergence…

Machine Learning · Computer Science 2026-04-10 Fabricio Macià , Shu Nakamura

Weight matrices in deep networks exhibit geometric continuity -- principal singular vectors of adjacent layers point in similar directions. While this property has been widely observed, its origin remains unexplained. Through experiments on…

Machine Learning · Computer Science 2026-05-07 Kyungwon Jeong , Won-Gi Paeng , Honggyo Suh

We consider the super-critical contact process on $\mathbb{Z}^d$. It is known that measures which dominate the upper invariant measure $\mu$ converge exponentially fast to $\mu$. However, the same is not true for measures which are below…

Probability · Mathematics 2013-10-24 Florian Völlering

The memorization capacity of neural networks with a given architecture has been thoroughly studied in many works. Specifically, it is well-known that memorizing $N$ samples can be done using a network of constant width, independent of $N$.…

Machine Learning · Computer Science 2025-02-18 Amitsour Egosi , Gilad Yehudai , Ohad Shamir

Understanding implicit bias of gradient descent for generalization capability of ReLU networks has been an important research topic in machine learning research. Unfortunately, even for a single ReLU neuron trained with the square loss, it…

Machine Learning · Computer Science 2022-06-14 Sangmin Lee , Byeongsu Sim , Jong Chul Ye

Gradient-only line searches (GOLS) adaptively determine step sizes along search directions for discontinuous loss functions resulting from dynamic mini-batch sub-sampling in neural network training. Step sizes in GOLS are determined by…

Machine Learning · Statistics 2020-02-25 D. Kafka , Daniel. N. Wilke

Recent approaches in the theoretical analysis of model-based deep learning architectures have studied the convergence of gradient descent in shallow ReLU networks that arise from generative models whose hidden layers are sparse. Motivated…

Machine Learning · Computer Science 2022-01-24 Emmanouil Theodosis , Bahareh Tolooshams , Pranay Tankala , Abiy Tasissa , Demba Ba

We formalize and interpret the geometric structure of $d$-dimensional fully connected ReLU layers in neural networks. The parameters of a ReLU layer induce a natural partition of the input domain, such that the ReLU layer can be…

Machine Learning · Computer Science 2023-11-09 Jonatan Vallin , Karl Larsson , Mats G. Larson

Gaussian Error Linear Unit (GELU) is a widely used smooth alternative to Rectifier Linear Unit (ReLU), yet many deployment, compression, and analysis toolchains are most naturally expressed for piecewise-linear (ReLU-type) networks. We…

In the Multi-Level Aggregation Problem (MLAP), requests arrive at the nodes of an edge-weighted tree T, and have to be served eventually. A service is defined as a subtree X of T that contains its root. This subtree X serves all requests…

We study the type of solutions to which stochastic gradient descent converges when used to train a single hidden-layer multivariate ReLU network with the quadratic loss. Our results are based on a dynamical stability analysis. In the…

Machine Learning · Computer Science 2023-07-03 Mor Shpigel Nacson , Rotem Mulayoff , Greg Ongie , Tomer Michaeli , Daniel Soudry

We introduce the blockwise gluing construction. This describes residuated integral chains which can be decomposed into (possibly) partial algebras, stacked one on top of the other, and such that elements in a certain component multiply in…

Logic · Mathematics 2025-12-22 Valeria Giustarini , Sara Ugolini

We study the implicit bias of gradient descent methods in solving a binary classification problem over a linearly separable dataset. The classifier is described by a nonlinear ReLU model and the objective function adopts the exponential…

Machine Learning · Computer Science 2018-10-17 Tengyu Xu , Yi Zhou , Kaiyi Ji , Yingbin Liang

Two networks are equivalent if they produce the same output for any given input. In this paper, we study the possibility of transforming a deep neural network to another network with a different number of units or layers, which can be…

Machine Learning · Computer Science 2019-05-29 Abhinav Kumar , Thiago Serra , Srikumar Ramalingam
‹ Prev 1 4 5 6 7 8 10 Next ›