English
Related papers

Related papers: Subcritical Signal Propagation at Initialization i…

200 papers

We present a detailed discussion of scalar wave propagation and light intensity transport in three dimensional random dielectric media with optical gain. The intrinsic length and time scales of such amplifying systems are studied and…

Disordered Systems and Neural Networks · Physics 2015-05-28 R. Frank , A. Lubatsch

To decrease the training overhead and improve the channel estimation accuracy in uplink cloud radio access networks (C-RANs), a superimposed-segment training design is proposed. The core idea of the proposal is that each mobile station…

Information Theory · Computer Science 2015-06-23 Xinqian Xie , Mugen Peng , H. Vincent Poor

Transformer models have achieved remarkable results in a wide range of applications. However, their scalability is hampered by the quadratic time and memory complexity of the self-attention mechanism concerning the sequence length. This…

Machine Learning · Computer Science 2024-02-27 Yury Nahshan , Joseph Kampeas , Emir Haleva

The relation between the jump length probability distribution function and the spectral line profile in resonance atomic radiation trapping is considered for Partial Frequency Redistribution (PFR) between absorbed and reemitted radiation.…

The quantized neural networks (QNNs) can be useful for neural network acceleration and compression, but during the training process they pose a challenge: how to propagate the gradient of loss function through the graph flow with a…

Machine Learning · Computer Science 2020-03-26 Jun Chen , Yong Liu , Hao Zhang , Shengnan Hou , Jian Yang

Autoregressive neural simulators now match classical solvers on short-horizon prediction of physical systems, yet their accuracy degrades rapidly when rolled out over long horizons. In this work, we identify transient amplification of…

Machine Learning · Computer Science 2026-05-18 Adeel Pervez , Francesco Locatello

Attention sinks and massive activations are recurring and closely related phenomena in Transformer models. Existing explanations have largely focused on the forward pass, yet in pre-norm Transformers, large residual-stream norms play only…

Machine Learning · Computer Science 2026-05-07 Yihong Chen , Zhouchen Lin , Quanming Yao

Regularization is typically understood as improving generalization by altering the landscape of local extrema to which the model eventually converges. Deep neural networks (DNNs), however, challenge this view: We show that removing…

Machine Learning · Computer Science 2019-06-03 Aditya Golatkar , Alessandro Achille , Stefano Soatto

We study the use of feedforward neural networks (FNN) to develop models of nonlinear dynamical systems from data. Emphasis is placed on predictions at long times, with limited data availability. Inspired by global stability analysis, and…

Machine Learning · Statistics 2020-06-16 Shaowu Pan , Karthik Duraisamy

Neural networks trained via gradient descent with random initialization and without any regularization enjoy good generalization performance in practice despite being highly overparametrized. A promising direction to explain this phenomenon…

Machine Learning · Computer Science 2022-05-17 Hancheng Min , Salma Tarmoun , Rene Vidal , Enrique Mallada

We study randomly initialized residual networks using mean field theory and the theory of difference equations. Classical feedforward neural networks, such as those with tanh activations, exhibit exponential behavior on the average when…

Neural and Evolutionary Computing · Computer Science 2017-12-27 Greg Yang , Samuel S. Schoenholz

Most graph-network-based meta-learning approaches model instance-level relation of examples. We extend this idea further to explicitly model the distribution-level relation of one example to all other examples in a 1-vs-N manner. We propose…

Computer Vision and Pattern Recognition · Computer Science 2020-04-02 Ling Yang , Liangliang Li , Zilun Zhang , Xinyu Zhou , Erjin Zhou , Yu Liu

Equilibrium propagation (EP) is a compelling alternative to the backpropagation of error algorithm (BP) for computing gradients of neural networks on biological or analog neuromorphic substrates. Still, the algorithm requires weight…

Machine Learning · Computer Science 2024-04-09 Axel Laborieux , Friedemann Zenke

The capacity of neural networks like the widely adopted transformer is known to be very high. Evidence is emerging that they learn successfully due to inductive bias in the training routine, typically a variant of gradient descent (GD). To…

Machine Learning · Computer Science 2023-03-09 William Merrill , Vivek Ramanujan , Yoav Goldberg , Roy Schwartz , Noah Smith

This paper analyses LightGCN in the context of graph recommendation algorithms. Despite the initial design of Graph Convolutional Networks for graph classification, the non-linear operations are not always essential. LightGCN enables linear…

Information Retrieval · Computer Science 2023-12-29 Milena Kapralova , Luca Pantea , Andrei Blahovici

Batch normalization (BN) is comprised of a normalization component followed by an affine transformation and has become essential for training deep neural networks. Standard initialization of each BN in a network sets the affine…

Computer Vision and Pattern Recognition · Computer Science 2022-07-18 Jim Davis , Logan Frank

Large language models are remarkably capable, yet how computation propagates through their layers remains poorly understood. A growing line of work treats depth as discrete time and the residual stream as a dynamical system, where each…

Machine Learning · Computer Science 2026-05-15 Jesseba Fernando , Grigori Guitchounts

Graph Convolutional Networks (GCNs) have been widely studied for compact data representation and semi-supervised learning tasks. However, existing GCNs usually use a fixed neighborhood graph which is not guaranteed to be optimal for…

Computer Vision and Pattern Recognition · Computer Science 2019-11-22 Bo Jiang , Leiling Wang , Jin Tang , Bin Luo

We analyze how the transient dynamics of large dynamical systems in the vicinity of a stationary point, modeled by a set of randomly coupled linear differential equations, depends on the network topology. We characterize the transient…

Adaptation and Self-Organizing Systems · Physics 2024-01-17 Wojciech Tarnowski , Izaak Neri , Pierpaolo Vivo

Joint network topology inference represents a canonical problem of jointly learning multiple graph Laplacian matrices from heterogeneous graph signals. In such a problem, a widely employed assumption is that of a simple common component…

Statistics Theory · Mathematics 2021-07-09 Yanli Yuan , De Wen Soh , Xiao Yang , Kun Guo , Tony Q. S. Quek
‹ Prev 1 3 4 5 6 7 10 Next ›