English
Related papers

Related papers: On the Properties of the Softmax Function with App…

200 papers

The paper is devoted to the development of new sufficient conditions for the calmness and the Aubin property of implicit multifunctions. As the basic tool one employs the directional limiting coderivative which, together with the graphical…

Optimization and Control · Mathematics 2016-11-28 Helmut Gfrerer , Jiří V. Outrata

In these notes, we present a general result concerning the Lipschitz regularity of a certain type of set-valued maps often found in constrained optimization and control problems. The class of multifunctions examined in this paper is…

Optimization and Control · Mathematics 2007-05-23 M. Papi , S. Sbaraglia

The Softmax function is ubiquitous in machine learning, multiple previous works suggested faster alternatives for it. In this paper we propose a way to compute classical Softmax with fewer memory accesses and hypothesize that this reduction…

Performance · Computer Science 2018-07-31 Maxim Milakov , Natalia Gimelshein

We establish risk bounds for Regularized Empirical Risk Minimizers (RERM) when the loss is Lipschitz and convex and the regularization function is a norm. In a first part, we obtain these results in the i.i.d. setup under subgaussian…

Statistics Theory · Mathematics 2021-01-07 Geoffrey Chinot , Guillaume Lecué , Matthieu Lerasle

We calculate the soft function using lattice QCD in the framework of large momentum effective theory incorporating the one-loop perturbative contributions. The soft function is a crucial ingredient in the lattice determination of light cone…

Softmax working with cross-entropy is widely used in classification, which evaluates the similarity between two discrete distribution columns (predictions and true labels). Inspired by chi-square test, we designed a new loss function called…

Machine Learning · Computer Science 2021-09-01 Zeyu Wang , Meiqing Wang

Softmax Loss (SL) is widely applied in recommender systems (RS) and has demonstrated effectiveness. This work analyzes SL from a pairwise perspective, revealing two significant limitations: 1) the relationship between SL and conventional…

Machine Learning · Computer Science 2025-08-05 Weiqin Yang , Jiawei Chen , Xin Xin , Sheng Zhou , Binbin Hu , Yan Feng , Chun Chen , Can Wang

We propose an accelerated meta-algorithm, which allows to obtain accelerated methods for convex unconstrained minimization in different settings. As an application of the general scheme we propose nearly optimal methods for minimizing…

Given a Lipschitz or smooth convex function $\, f:K \to \mathbb{R}$ for a bounded polytope $K \subseteq \mathbb{R}^d$ defined by $m$ inequalities, we consider the problem of sampling from the log-concave distribution $\pi(\theta) \propto…

Data Structures and Algorithms · Computer Science 2022-11-16 Oren Mangoubi , Nisheeth K. Vishnoi

The convexification numerical method with the rigorously established global convergence property is constructed for a problem for the Mean Field Games System of the second order. This is the problem of the retrospective analysis of a game…

Numerical Analysis · Mathematics 2023-06-30 Michael V. Klibanov , Jingzhi Li , Zhipeng Yang

We introduce a notion of inexact model of a convex objective function, which allows for errors both in the function and in its gradient. For this situation, a gradient method with an adaptive adjustment of some parameters of the model is…

Optimization and Control · Mathematics 2021-10-12 Fedor S. Stonyakin

In reinforcement learning (RL), different reward functions can define the same optimal policy but result in drastically different learning performance. For some, the agent gets stuck with a suboptimal behavior, and for others, it solves the…

Machine Learning · Computer Science 2025-02-25 Grigorii Veviurko , Wendelin Böhmer , Mathijs de Weerdt

Preconditioning is a crucial operation in gradient-based numerical optimisation. It helps decrease the local condition number of a function by appropriately transforming its gradient. For a convex function, where the gradient can be…

Optimization and Control · Mathematics 2023-08-29 Dmitrii A. Pasechnyuk , Alexander Gasnikov , Martin Takáč

Achieving convergence of multiple learning agents in general $N$-player games is imperative for the development of safe and reliable machine learning (ML) algorithms and their application to autonomous systems. Yet it is known that, outside…

Computer Science and Game Theory · Computer Science 2023-01-24 Aamal Abbas Hussain , Francesco Belardinelli , Georgios Piliouras

In the context of structured nonconvex optimization, we estimate the increase in minimum value for a decision that is robust to parameter perturbations as compared to the value of a nominal problem. The estimates rely on detailed…

Optimization and Control · Mathematics 2022-11-22 Johannes O. Royset

The standard assumption for proving linear convergence of first order methods for smooth convex optimization is the strong convexity of the objective function, an assumption which does not hold for many practical applications. In this…

Optimization and Control · Mathematics 2016-08-10 I. Necoara , Yu. Nesterov , F. Glineur

In this paper, we study inverse game theory (resp. inverse multiagent learning) in which the goal is to find parameters of a game's payoff functions for which the expected (resp. sampled) behavior is an equilibrium. We formulate these…

Computer Science and Game Theory · Computer Science 2025-02-21 Denizalp Goktas , Amy Greenwald , Sadie Zhao , Alec Koppel , Sumitra Ganesh

In this work we deal with the stochastic homogenization of the initial boundary value problems of monotone type. The models of monotone type under consideration describe the deformation behaviour of inelastic materials with a microstructure…

Analysis of PDEs · Mathematics 2017-01-16 Martin Heida , Sergiy Nesenenko

Adaptive momentum methods have recently attracted a lot of attention for training of deep neural networks. They use an exponential moving average of past gradients of the objective function to update both search directions and learning…

Optimization and Control · Mathematics 2021-04-27 Babak Barazandeh , Davoud Ataee Tarzanagh , George Michailidis

The task of approximating an arbitrary convex function arises in several learning problems such as convex regression, learning with a difference of convex (DC) functions, and learning Bregman or $f$-divergences. In this paper, we develop…