English
Related papers

Related papers: Target Entropy Annealing for Discrete Soft Actor-C…

200 papers

In this paper, we study the role of the critic in actor--critic for entropy-regularized, finite, discounted environments. We establish that, when the critic is exact, using the latter as a baseline is a variance-reduction method in a strong…

Machine Learning · Computer Science 2026-05-26 Safwan Labbi , Paul Mangold , Daniil Tiapkin , Eric Moulines

Maximum entropy has become a mainstream off-policy reinforcement learning (RL) framework for balancing exploitation and exploration. However, two bottlenecks still limit further performance improvement: (1) non-stationary Q-value estimation…

Machine Learning · Computer Science 2025-11-18 Guojian Zhan , Likun Wang , Pengcheng Wang , Feihong Zhang , Jingliang Duan , Masayoshi Tomizuka , Shengbo Eben Li

In this work, we propose Behavior-Guided Actor-Critic (BAC), an off-policy actor-critic deep RL algorithm. BAC mathematically formulates the behavior of the policy through autoencoders by providing an accurate estimation of how frequently…

Machine Learning · Computer Science 2021-04-12 Ammar Fayad , Majd Ibrahim

Iterative generative policies, such as diffusion models and flow matching, offer superior expressivity for continuous control but complicate Maximum Entropy Reinforcement Learning because their action log-densities are not directly…

Machine Learning · Computer Science 2026-02-16 Lei Lv , Yunfei Li , Yu Luo , Fuchun Sun , Xiao Ma

Spatially selective active noise control (SSANC) hearables aim to attenuate noise from certain directions at the eardrum while preserving desired speech arriving from selected directions. Existing SSANC systems typically assume an accurate…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Tong Xiao , Reinhild Roden , Matthias Blau , Simon Doclo

Automated grading systems can efficiently score short-answer responses, yet they often fail to indicate when a grading decision is uncertain or potentially contentious. We introduce semantic entropy, a measure of variability across multiple…

Artificial Intelligence · Computer Science 2025-08-07 Karrtik Iyer , Manikandan Ravikiran , Prasanna Pendse , Shayan Mohanty

In this paper, the Entropically Damped Artificial Compressibility (EDAC) formulation of Clausen (2013) is used in the context of the Smoothed Particle Hydrodynamics (SPH) method for the simulation of incompressible fluids. Traditionally,…

Computational Physics · Physics 2018-11-30 Prabhu Ramachandran , Kunal Puri

Maximum Entropy Reinforcement Learning (MaxEnt RL) algorithms such as Soft Q-Learning (SQL) and Soft Actor-Critic trade off reward and policy entropy, which has the potential to improve training stability and robustness. Most MaxEnt RL…

Machine Learning · Computer Science 2021-11-30 Dailin Hu , Pieter Abbeel , Roy Fox

The deployment of robots in uncontrolled environments requires them to operate robustly under previously unseen scenarios, like irregular terrain and wind conditions. Unfortunately, while rigorous safety frameworks from robust optimal…

Machine Learning · Computer Science 2024-06-11 Kai-Chieh Hsu , Duy Phuong Nguyen , Jaime Fernández Fisac

High-precision control tasks present substantial challenges for reinforcement learning (RL) algorithms, frequently resulting in suboptimal performance attributed to network approximation inaccuracies and inadequate sample quality.These…

Machine Learning · Computer Science 2025-02-05 Donghe Chen , Yubin Peng , Tengjie Zheng , Han Wang , Chaoran Qu , Lin Cheng

Recent advances in spatially selective active noise control (SSANC) using multiple microphones have enabled hearables to suppress undesired noise while preserving desired speech from a specific direction. Aiming to achieve minimal speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-18 Tong Xiao , Reinhild Roden , Matthias Blau , Simon Doclo

In this paper, we study the problem of multiple stochastic agents interacting in a dynamic game scenario with continuous state and action spaces. We define a new notion of stochastic Nash equilibrium for boundedly rational agents, which we…

Optimization and Control · Mathematics 2021-10-05 Negar Mehr , Mingyu Wang , Mac Schwager

Maximum entropy (MaxEnt) RL maximizes a combination of the original task reward and an entropy reward. It is believed that the regularization imposed by entropy, on both policy improvement and policy evaluation, together contributes to good…

Machine Learning · Computer Science 2022-02-01 Haonan Yu , Haichao Zhang , Wei Xu

Temporal Action Localization (TAL) methods typically operate on top of feature sequences from a frozen snippet encoder that is pretrained with the Trimmed Action Classification (TAC) tasks, resulting in a task discrepancy problem. While…

Computer Vision and Pattern Recognition · Computer Science 2022-11-14 Hyolim Kang , Hanjung Kim , Joungbin An , Minsu Cho , Seon Joo Kim

Deciding what and when to observe is critical when making observations is costly. In a medical setting where observations can be made sequentially, making these observations (or not) should be an active choice. We refer to this as the…

Machine Learning · Computer Science 2019-06-18 Jinsung Yoon , James Jordon , Mihaela van der Schaar

Traditional Reinforcement Learning (RL) policies are typically implemented with fixed control rates, often disregarding the impact of control rate selection. This can lead to inefficiencies as the optimal control rate varies with task…

Robotics · Computer Science 2024-08-13 Dong Wang , Giovanni Beltrame

We analyze the global convergence of the single-timescale actor-critic (AC) algorithm for the infinite-horizon discounted Markov Decision Processes (MDPs) with finite state spaces. To this end, we introduce an elegant analytical framework…

Machine Learning · Computer Science 2025-06-05 Navdeep Kumar , Priyank Agrawal , Giorgia Ramponi , Kfir Yehuda Levy , Shie Mannor

Expressive generative policies such as diffusion and flow models are appealing for MaxEnt online reinforcement learning because of their ability to model multimodal and highly non-Gaussian action distributions. However, training effective…

Machine Learning · Computer Science 2026-05-12 Ke He , Le He , Shunpu Tang , Yafei Wang , Lisheng Fan

Entropy regularization is an important idea in reinforcement learning, with great success in recent algorithms like Soft Q Network (SQN) and Soft Actor-Critic (SAC1). In this work, we extend this idea into the on-policy realm. We propose…

Machine Learning · Computer Science 2020-10-19 Jingbin Liu , Xinyang Gu , Shuai Liu

Generating competitive strategies and performing continuous motion planning simultaneously in an adversarial setting is a challenging problem. In addition, understanding the intent of other agents is crucial to deploying autonomous systems…

Robotics · Computer Science 2025-06-17 Hongrui Zheng , Zhijun Zhuang , Stephanie Wu , Shuo Yang , Rahul Mangharam