English
Related papers

Related papers: Transformative or Conservative? Conservation laws …

200 papers

We propose and analyze a new family of algorithms for training neural networks with ReLU activations. Our algorithms are based on the technique of alternating minimization: estimating the activation patterns of each ReLU for all given…

Machine Learning · Computer Science 2018-10-12 Gauri Jagatap , Chinmay Hegde

The loss surface of deep neural networks has recently attracted interest in the optimization and machine learning communities as a prime example of high-dimensional non-convex problem. Some insights were recently gained using spin glass…

Machine Learning · Statistics 2017-06-05 C. Daniel Freeman , Joan Bruna

In this concise contribution, it is demonstrated that Recurrent Neural Networks (RNNs) based on Gated Recurrent Unit (GRU) architecture, possess the capability to learn the complex dynamics of rate-and-state friction (RSF) laws from…

Geophysics · Physics 2025-05-22 Gaëtan Cortes , Joaquin Garcia-Suarez

The relationship between the properties of a dynamical system and the structure of its defining equations has long been studied in many contexts. Here we study this problem for the class of conjunctive (resp. disjunctive) Boolean networks,…

Combinatorics · Mathematics 2008-05-13 Abdul Salam Jarrah , Reinhard Laubenbacher , Alan Veliz-Cuba

In this work, we propose Retentive Network (RetNet) as a foundation architecture for large language models, simultaneously achieving training parallelism, low-cost inference, and good performance. We theoretically derive the connection…

Computation and Language · Computer Science 2023-08-10 Yutao Sun , Li Dong , Shaohan Huang , Shuming Ma , Yuqing Xia , Jilong Xue , Jianyong Wang , Furu Wei

Hyperbolic conservation laws govern a wide range of transport-driven dynamics featuring shocks, contact discontinuities, and complex wave interactions, posing distinct challenges for deep-learning-based surrogate modeling. While classical…

Computational Physics · Physics 2026-04-20 Jiamin Jiang , Shanglin Lv , Jingrun Chen

We investigate what can be learned from translating numerical algorithms into neural networks. On the numerical side, we consider explicit, accelerated explicit, and implicit schemes for a general higher order nonlinear diffusion equation…

Numerical Analysis · Mathematics 2021-05-18 Tobias Alt , Pascal Peter , Joachim Weickert , Karl Schrader

This work presents an analysis of the effectiveness of using standard shallow feed-forward networks to mimic the behavior of the attention mechanism in the original Transformer model, a state-of-the-art architecture for sequence-to-sequence…

Computation and Language · Computer Science 2024-02-06 Vukasin Bozic , Danilo Dordevic , Daniele Coppola , Joseph Thommes , Sidak Pal Singh

Data-centric methods have shown great potential in understanding and predicting spatiotemporal dynamics, enabling better design and control of the object system. However, deep learning models often lack interpretability, fail to obey…

Machine Learning · Computer Science 2025-01-07 Yuan Mi , Pu Ren , Hongteng Xu , Hongsheng Liu , Zidong Wang , Yike Guo , Ji-Rong Wen , Hao Sun , Yang Liu

Gradients of neural networks encode valuable information for optimization, editing, and analysis of models. Therefore, practitioners often treat gradients as inputs to task-specific algorithms, e.g. for pruning or optimization. Recent works…

Machine Learning · Computer Science 2025-10-14 Yoav Gelberg , Yam Eitan , Aviv Navon , Aviv Shamsian , Theo , Putterman , Michael Bronstein , Haggai Maron

In this paper we address the importance and the impact of employing structure preserving neural networks as surrogate of the analytical physics-based models typically employed to describe the rheology of non-Newtonian fluids in Stokes…

Numerical Analysis · Mathematics 2024-01-17 Nicola Parolini , Andrea Poiatti , Julian Vene' , Marco Verani

The system of equations of one-dimensional shallow water over uneven bottom in Euler's and Lagrange's variables is considered. Intermediate system of equations is introduced. Hydrodynamic conservation laws of intermediate system of…

Exactly Solvable and Integrable Systems · Physics 2018-12-14 Alexander V. Aksenov , Konstantin P. Druzhkov

We study the convergence of gradient flow for the training of deep neural networks. If Residual Neural Networks are a popular example of very deep architectures, their training constitutes a challenging optimization problem due notably to…

Machine Learning · Computer Science 2025-07-22 Raphaël Barboni , Gabriel Peyré , François-Xavier Vialard

Convolutional neural networks (CNNs) often perform well, but their stability is poorly understood. To address this problem, we consider the simple prototypical problem of signal denoising, where classical approaches such as nonlinear…

Machine Learning · Computer Science 2020-06-09 Tobias Alt , Joachim Weickert , Pascal Peter

Radiation hydrodynamics simulations based on the one-fluid two-temperature model may violate the law of energy conservation because the governing equations are expressed in a nonconservative formulation. Here, we maintain the important…

Plasma Physics · Physics 2018-03-09 Takashi Shiroto , Soshi Kawai , Naofumi Ohnishi

In this paper, we develop novel techniques that can be used to alter the architecture of a neural network, while maintaining the function it represents. Such operations are known as function preserving transforms and have proven useful in…

Machine Learning · Computer Science 2024-10-16 Michael Painter

This report analyzes the evolution of key design patterns in computer vision by examining six influential papers. The analysis begins with foundational architectures for image recognition. We review ResNet, which introduced residual…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Radu-Andrei Bourceanu , Neil De La Fuente , Jan Grimm , Andrei Jardan , Andriy Manucharyan , Cornelius Weiss , Daniel Cremers , Roman Pflugfelder

We propose ReDense as a simple and low complexity way to improve the performance of trained neural networks. We use a combination of random weights and rectified linear unit (ReLU) activation function to add a ReLU dense (ReDense) layer to…

Machine Learning · Computer Science 2020-10-27 Alireza M. Javid , Sandipan Das , Mikael Skoglund , Saikat Chatterjee

Recent approaches in the theoretical analysis of model-based deep learning architectures have studied the convergence of gradient descent in shallow ReLU networks that arise from generative models whose hidden layers are sparse. Motivated…

Machine Learning · Computer Science 2022-01-24 Emmanouil Theodosis , Bahareh Tolooshams , Pranay Tankala , Abiy Tasissa , Demba Ba

Generative Flow Networks (GFlowNets) learn to sample states proportional to an unnormalized reward. Despite their theoretical promise, practical training is often unstable, exhibiting severe loss spikes and mode collapse. To tackle this, we…

‹ Prev 1 8 9 10 Next ›