English
Related papers

Related papers: On the Properties of the Softmax Function with App…

200 papers

The usual approach to developing and analyzing first-order methods for smooth convex optimization assumes that the gradient of the objective function is uniformly smooth with some Lipschitz constant $L$. However, in many settings the…

Optimization and Control · Mathematics 2017-10-11 Haihao Lu , Robert M. Freund , Yurii Nesterov

We introduce and analyze an algorithm for the minimization of convex functions that are the sum of differentiable terms and proximable terms composed with linear operators. The method builds upon the recently developed smoothed gap…

Optimization and Control · Mathematics 2017-06-20 Quang Van Nguyen , Olivier Fercoq , Volkan Cevher

We study differentiability properties of convex operators defined on a Banach space with values in an $\Lc_p$ space and of their compositions with monotonic convex functionals on this space. We develop new tools for operators enjoying an…

Optimization and Control · Mathematics 2025-11-10 Darinka Dentcheva , Andrzej Ruszczynski

The structural properties of graphs are usually characterized in terms of invariants, which are functions of graphs that do not depend on the labeling of the nodes. In this paper we study convex graph invariants, which are graph invariants…

Optimization and Control · Mathematics 2012-09-21 Venkat Chandrasekaran , Pablo A. Parrilo , Alan S. Willsky

The familiar second derivative test for convexity, combined with resolvent calculus, is shown to yield a useful tool for the study of convex matrix-valued functions. We demonstrate the applicability of this approach on a number of theorems…

Quantum Physics · Physics 2024-07-26 Michael Aizenman , Giorgio Cipolloni

Researchers are increasingly focusing on intelligent games as a hot research area.The article proposes an algorithm that combines the multi-attribute management and reinforcement learning methods, and that combined their effect on…

Artificial Intelligence · Computer Science 2021-09-07 Yuxiang Sun , Bo Yuan , Yufan Xue , Jiawei Zhou , Xiaoyu Zhang , Xianzhong Zhou

In the present paper, the following convexity principle is proved: any closed convex multifunction, which is metrically regular in a certain uniform sense near a given point, carries small balls centered at that point to convex sets, even…

Optimization and Control · Mathematics 2015-04-13 Amos Uderzo

We consider the problem of global optimization of an unknown non-convex smooth function with zeroth-order feedback. In this setup, an algorithm is allowed to adaptively query the underlying function at different locations and receives noisy…

Machine Learning · Statistics 2018-03-26 Yining Wang , Sivaraman Balakrishnan , Aarti Singh

Typically, Softmax is used in the final layer of a neural network to get a probability distribution for output classes. But the main problem with Softmax is that it is computationally expensive for large scale data sets with large number of…

Machine Learning · Computer Science 2018-12-17 Abdul Arfat Mohammed , Venkatesh Umaashankar

Deep networks have enabled reinforcement learning to scale to more complex and challenging domains, but these methods typically require large quantities of training data. An alternative is to use sample-efficient episodic control methods:…

Machine Learning · Computer Science 2019-11-22 Marta Sarrico , Kai Arulkumaran , Andrea Agostinelli , Pierre Richemond , Anil Anthony Bharath

The monotonicity of the Mittag-Leffler function $E_{\alpha}$ with respect to the parameter $\alpha$ is investigated, via some convex ordering properties for related random variables. In particular, it is shown that the mapping…

Classical Analysis and ODEs · Mathematics 2025-12-05 Rui Ferreira , Thomas Simon

Exploration-exploitation dilemma has long been a crucial issue in reinforcement learning. In this paper, we propose a new approach to automatically balance between these two. Our method is built upon the Soft Actor-Critic (SAC) algorithm,…

Machine Learning · Computer Science 2020-08-03 Yufei Wang , Tianwei Ni

The Bellman equation and its continuous form, the Hamilton-Jacobi-Bellman equation, are ubiquitous in reinforcement learning and control theory. However, these equations become intractable for high-dimensional or nonlinear systems. This…

Artificial Intelligence · Computer Science 2026-05-04 Preston Rozwood , Edward Mehrez , Ludger Paehler , Wen Sun , Steven L. Brunton

This paper considers the problem of inverse reinforcement learning in zero-sum stochastic games when expert demonstrations are known to be not optimal. Compared to previous works that decouple agents in the game by assuming optimality in…

Machine Learning · Statistics 2018-06-07 Xingyu Wang , Diego Klabjan

We consider the problem of analyzing and designing gradient-based discrete-time optimization algorithms for a class of unconstrained optimization problems having strongly convex objective functions with Lipschitz continuous gradient. By…

Optimization and Control · Mathematics 2025-10-20 Simon Michalowsky , Carsten Scherer , Christian Ebenbauer

In this paper, we address the inverse problem for linear-quadratic differential non-cooperative games with output-feedback. Given players' stabilizing feedback laws, the goal is to find cost function parameters that lead to a game for which…

Optimization and Control · Mathematics 2024-10-27 Emin Martirosyan , Ming Cao

We consider the inverse problem of dynamic games, where cost function parameters are sought which explain observed behavior of interacting players. Maximum entropy inverse reinforcement learning is extended to the N-player case in order to…

Systems and Control · Electrical Eng. & Systems 2020-07-27 Jairo Inga , Esther Bischoff , Florian Köpf , Sören Hohmann

We show that the thermal subadditivity of entropy provides a common basis to derive a strong form of the bounded difference inequality and related results as well as more recent inequalities applicable to convex Lipschitz functions, random…

Statistics Theory · Mathematics 2012-05-09 Andreas Maurer

Penalty functions are widely used to enforce constraints in optimization problems and reinforcement leaning algorithms. Softplus and algebraic penalty functions are proposed to overcome the sensitivity of the Courant-Beltrami method to…

Optimization and Control · Mathematics 2021-07-12 Stefan Meili

Numerous models for supervised and reinforcement learning benefit from combinations of discrete and continuous model components. End-to-end learnable discrete-continuous models are compositional, tend to generalize better, and are more…

Machine Learning · Computer Science 2023-07-27 David Friede , Mathias Niepert