English
Related papers

Related papers: Linear $\beta$-reduction

200 papers

The Adam optimizer is a cornerstone of modern deep learning, yet the empirical necessity of each of its individual components is often taken for granted. This paper presents a focused investigation into the role of bias-correction, a…

Machine Learning · Computer Science 2025-11-27 Sam Laing , Antonio Orvieto

The supervised training of a deep neural network on a given dataset consists in the unconstrained minimization of the finite sum of continuously differentiable functions, commonly referred to as loss with respect to the samples. These…

Optimization and Control · Mathematics 2024-05-24 Corrado Coppola , Giampaolo Liuzzi , Laura Palagi

Estimation in generalized linear models (GLM) is complicated by the presence of constraints. One can handle constraints by maximizing a penalized log-likelihood. Penalties such as the lasso are effective in high dimensions, but often lead…

Machine Learning · Statistics 2017-11-07 Jason Xu , Eric C. Chi , Kenneth Lange

Deterministic 2-head finite automata which are machines that process an input word from both ends are analyzed for their ability to perform reversible computations. This implies that the automata are backward deterministic, enabling unique…

Formal Languages and Automata Theory · Computer Science 2025-07-22 Benedek Nagy , Walaa Yasin

Linear Discriminant Analysis (LDA) is a well-known method for dimensionality reduction and classification. Previous studies have also extended the binary-class case into multi-classes. However, many applications, such as object detection…

Machine Learning · Computer Science 2013-09-24 Gang Chen

Relative temporal-difference (TD) learning was introduced to mitigate the slow convergence of TD methods when the discount factor approaches one by subtracting a baseline from the temporal-difference update. While this idea has been studied…

Machine Learning · Computer Science 2026-04-08 Masoud S. Sakha , Rushikesh Kamalapurkar , Sean Meyn

This paper describes the notion of \sigma -symmetry, which extends the one of \lambda-symmetry, and its application to reduction procedures of systems of ordinary differential equations and of dynamical systems as well. We also consider…

Mathematical Physics · Physics 2015-06-16 Giampaolo Cicogna

Graph coarsening aims to diminish the size of a graph to lighten its memory footprint, and has numerous applications in graph signal processing and machine learning. It is usually defined using a reduction matrix and a lifting matrix,…

Machine Learning · Computer Science 2026-01-29 Antonin Joly , Nicolas Keriven , Aline Roumy

Optimizing Neural networks is a difficult task which is still not well understood. On the other hand, fixed representation methods such as kernels and random features have provable optimization guarantees but inferior performance due to…

Machine Learning · Computer Science 2024-01-17 Amit Daniely , Mariano Schain , Gilad Yehudai

Recovering the digital input of a time-discrete linear system from its (noisy) output is a significant challenge in the fields of data transmission, deconvolution, channel equalization, and inverse modeling. A variety of algorithms have…

Optimization and Control · Mathematics 2020-12-03 Sophie M. Fosson

Higher-order beta-matching is the following decision problem: given two simply typed lambda-terms, can the first term be instantiated to be beta-equivalent to the second term? This problem was formulated by Huet in the 1970s and shown…

Logic in Computer Science · Computer Science 2026-02-03 Andrej Dudenhefner

These notes are issued from a short course given by the author in a summer school in Chamb{\'e}ry in June 2015. We consider general semilinear PDE's and we address the following two questions: 1) How to design an efficient feedback control…

Analysis of PDEs · Mathematics 2015-06-22 Emmanuel Trélat

The beta process has recently been widely used as a nonparametric prior for different models in machine learning, including latent feature models. In this paper, we prove the asymptotic consistency of the finite dimensional approximation of…

Statistics Theory · Mathematics 2014-11-14 Luai Al Labadi , Mahmoud Zarepour

The convex transform order is one way to make precise comparison between the skewness of probability distributions on the real line. We establish a simple and complete characterisation of when one Beta distribution is smaller than another…

Probability · Mathematics 2021-01-01 Idir Arab , Paulo Eduardo Oliveira , Tilo Wiklund

We prove that given two cut free nets of linear logic, by means of their relational interpretations one can: 1) first determine whether or not the net obtained by cutting the two nets is strongly normalizable 2) then (in case it is strongly…

Logic in Computer Science · Computer Science 2014-08-28 Daniel de Carvalho , Lorenzo Tortora de Falco

Low rank regularization, in essence, involves introducing a low rank or approximately low rank assumption for matrix we aim to learn, which has achieved great success in many fields including machine learning, data mining and computer…

Computer Vision and Pattern Recognition · Computer Science 2020-12-11 Zhanxuan Hu , Feiping Nie , Rong Wang , Xuelong Li

ReRecent studies in machine learning are based on models in which parameters or state variables are bounded restricted. These restrictions are from prior information to ensure the validity of scientific theories or structural consistency…

Methodology · Statistics 2024-01-26 Solmaz Seifollahi , Hossein Bevrani , Kristofer Mansson

Optimisers are an essential component for training machine learning models, and their design influences learning speed and generalisation. Several studies have attempted to learn more effective gradient-descent optimisers via solving a…

Machine Learning · Computer Science 2022-03-08 Boyan Gao , Henry Gouk , Hae Beom Lee , Timothy M. Hospedales

Recently, it has been argued that encoder-decoder models can be made more interpretable by replacing the softmax function in the attention with its sparse variants. In this work, we introduce a novel, simple method for achieving sparsity in…

Computation and Language · Computer Science 2021-10-07 Biao Zhang , Ivan Titov , Rico Sennrich

Classical convergence theory of Runge-Kutta methods assumes that the time step is small relative to the Lipschitz constant of the ordinary differential equation (ODE). For stiff problems, that assumption is often violated, and a problematic…

Numerical Analysis · Mathematics 2026-05-05 Steven B. Roberts , David Shirokoff , Abhijit Biswas , Benjamin Seibold