English
Related papers

Related papers: Entropy-regularized penalization schemes and refle…

200 papers

This paper discusses the application of L1-regularized maximum entropy modeling or SL1-Max [9] to multiclass categorization problems. A new modification to the SL1-Max fast sequential learning algorithm is proposed to handle conditional…

Machine Learning · Computer Science 2007-05-23 Patrick Haffner , Steven Phillips , Rob Schapire

Entropic regularization of policies in Reinforcement Learning (RL) is a commonly used heuristic to ensure that the learned policy explores the state-space sufficiently before overfitting to a local optimal policy. The primary motivation for…

Machine Learning · Computer Science 2021-01-19 Hisham Husain , Kamil Ciosek , Ryota Tomioka

This article introduces an imitation learning method for learning maximum entropy policies that comply with constraints demonstrated by expert trajectories executing a task. The formulation of the method takes advantage of results…

Machine Learning · Computer Science 2025-07-10 George Papadopoulos , George A. Vouros

In this paper, we study reflected backward stochastic differential equation (reflected BSDE in abbreviation) with rank-based data in a Markovian framework; that is, the solution to the reflected BSDE is above a prescribed boundary process…

Probability · Mathematics 2020-07-14 Zhen-Qing Chen , Xinwei Feng

In this paper, we study existence and uniqueness to multidimensional Reflected Backward Stochastic Differential Equation in an open convex domain, allowing for oblique directions of reflection. In a Markovian framework, combining \emph{a…

Probability · Mathematics 2018-07-18 Jean-François Chassagneux , Adrien Richou

This paper studies the continuous-time reinforcement learning for stochastic singular control with the application to an infinite-horizon irreversible reinsurance problems. The singular control is equivalently characterized as a pair of…

Optimization and Control · Mathematics 2025-12-03 Zongxia Liang , Xiaodong Luo , Xiang Yu

In this paper we study, by probabilistic techniques, the convergence of the value function for a two-scale, infinite-dimensional, stochastic controlled system as the ratio between the two evolution speeds diverges. The value function is…

Optimization and Control · Mathematics 2018-09-12 Giuseppina Guatteri , Gianmario Tessitore

One of the most critical challenges in deep reinforcement learning is to maintain the long-term exploration capability of the agent. To tackle this problem, it has been recently proposed to provide intrinsic rewards for the agent to…

Machine Learning · Computer Science 2022-06-02 Mingqi Yuan , Man-on Pun , Dong Wang

In this paper, we provide two new stable online algorithms for the problem of prediction in reinforcement learning, \emph{i.e.}, estimating the value function of a model-free Markov reward process using the linear function approximation…

Machine Learning · Computer Science 2018-06-19 Ajin George Joseph , Shalabh Bhatnagar

Simplicity is a critical inductive bias for designing data-driven controllers, especially when robustness is important. Despite the impressive results of deep reinforcement learning in complex control tasks, it is prone to capturing…

Machine Learning · Computer Science 2025-05-09 Bang You , Chenxu Wang , Huaping Liu

In this paper, we investigate the well-posedness of bounded and unbounded solutions for reflected backward stochastic differential equations (RBSDEs) and backward stochastic differential equations (BSDEs). The generators of these equations…

Probability · Mathematics 2026-04-21 Shiqiu Zheng

We consider the representation of the value of a class of optimal stopping problems of linear diffusions in a linearized form as an expected supremum of a known function. We establish an explicit integral representation of this representing…

Probability · Mathematics 2017-03-16 Luis H. R. Alvarez E. , Pekka Matomäki

I formulate an entropy-rate maximization problem at the observable level for stochastic processes observed through an information-reducing observation map. For a visible stationary law, the map determines an observational fiber of hidden…

Information Theory · Computer Science 2026-04-14 Oleg Kiriukhin

In this paper we consider general rank minimization problems with rank appearing in either objective function or constraint. We first establish that a class of special rank minimization problems has closed-form solutions. Using this result,…

Optimization and Control · Mathematics 2012-05-30 Zhaosong Lu , Yong Zhang

We present a method of automatically synthesizing steps to solve search problems. Given a specification of a search problem, our approach uses symbolic execution to analyze the specification in order to extract a set of constraints which…

Logic in Computer Science · Computer Science 2020-09-24 Mara Downing , Abtin Molavi , Lucas Bang

In this note we propose a new approach towards solving numerically optimal stopping problems via reinforced regression based Monte Carlo algorithms. The main idea of the method is to reinforce standard linear regression algorithms in each…

Numerical Analysis · Mathematics 2019-07-02 Denis Belomestny , John Schoenmakers , Vladimir Spokoiny , Bakhyt Zharkynbay

We show existence of a unique solution and a comparison theorem for a one-dimensional backward stochastic differential equation with jumps that emerge from a L\'evy process. The considered generators obey a time-dependent extended…

Probability · Mathematics 2019-01-21 Christel Geiss , Alexander Steinicke

Entropy regularized algorithms such as Soft Q-learning and Soft Actor-Critic, recently showed state-of-the-art performance on a number of challenging reinforcement learning (RL) tasks. The regularized formulation modifies the standard RL…

Machine Learning · Statistics 2019-10-15 Elena Smirnova , Elvis Dohmatob

In this paper we make a survey on the so called randomization method, a recent methodology to study stochastic optimization problems. It allows to represent the value function of an optimal control problem by a suitable backward stochastic…

Optimization and Control · Mathematics 2025-06-12 Marco Fuhrman

An effective way to scale up test-time compute of large language models is to sample multiple responses and then select the best one, as in Grok Heavy and Gemini Deep Think. Existing selection methods often rely on external reward models,…

Machine Learning · Computer Science 2026-05-04 Wenshuo Zhao , Qi Zhu , Xingshan Zeng , Fei Mi , Lifeng Shang , Yi R. , Fung
‹ Prev 1 8 9 10 Next ›