中文
相关论文

相关论文: 6D (2,0) Bootstrap with soft-Actor-Critic

200 篇论文

Expressive generative policies such as diffusion and flow models are appealing for MaxEnt online reinforcement learning because of their ability to model multimodal and highly non-Gaussian action distributions. However, training effective…

机器学习 · 计算机科学 2026-05-12 Ke He , Le He , Shunpu Tang , Yafei Wang , Lisheng Fan

This work extends an established critic match loss landscape visualization method from online to off-policy reinforcement learning (RL), aiming to reveal the optimization geometry behind critic learning. Off-policy RL differs from stepwise…

机器学习 · 计算机科学 2026-03-17 Jingyi Liu , Jian Guo , Eberhard Gill

We aim to develop off-policy DRL algorithms that not only exceed state-of-the-art performance but are also simple and minimalistic. For standard continuous control benchmarks, Soft Actor-Critic (SAC), which employs entropy maximization,…

机器学习 · 计算机科学 2020-12-08 Che Wang , Yanqiu Wu , Quan Vuong , Keith Ross

Actor-critic methods, a type of model-free Reinforcement Learning, have been successfully applied to challenging tasks in continuous control, often achieving state-of-the art performance. However, wide-scale adoption of these methods in…

机器学习 · 统计学 2019-10-29 Kamil Ciosek , Quan Vuong , Robert Loftin , Katja Hofmann

Low-precision training has become a popular approach to reduce compute requirements, memory footprint, and energy consumption in supervised learning. In contrast, this promising approach has not yet enjoyed similarly widespread adoption…

机器学习 · 计算机科学 2021-06-07 Johan Bjorck , Xiangyu Chen , Christopher De Sa , Carla P. Gomes , Kilian Q. Weinberger

Decentralized Actor-Critic (AC) algorithms have been widely utilized for multi-agent reinforcement learning (MARL) and have achieved remarkable success. Apart from its empirical success, the theoretical convergence property of decentralized…

机器学习 · 计算机科学 2023-01-31 Qijun Luo , Xiao Li

The soft bootstrap is an on-shell method to constrain the landscape of effective field theories (EFTs) of massless particles via the consistency of the low-energy S-matrix. Given assumptions on the on-shell data (particle spectra, linear…

高能物理 - 理论 · 物理学 2019-02-20 Henriette Elvang , Marios Hadjiantonis , Callum R. T. Jones , Shruti Paranjape

Motivated by applications to critical phenomena and open theoretical questions, we study conformal field theories with $O(m)\times O(n)$ global symmetry in $d=3$ spacetime dimensions. We use both analytic and numerical bootstrap techniques.…

高能物理 - 理论 · 物理学 2020-12-10 Johan Henriksson , Stefanos R. Kousvos , Andreas Stergiou

We analyze the constraints imposed by unitarity and crossing symmetry on the four-point function of the stress-tensor multiplet of ${\cal N}=8$ superconformal field theories in three dimensions. We first derive the superconformal blocks by…

高能物理 - 理论 · 物理学 2015-06-22 Shai M. Chester , Jaehoon Lee , Silviu S. Pufu , Ran Yacoby

Deep reinforcement learning (RL) algorithms can use high-capacity deep networks to learn directly from image observations. However, these high-dimensional observation spaces present a number of challenges in practice, since the policy must…

机器学习 · 计算机科学 2020-10-27 Alex X. Lee , Anusha Nagabandi , Pieter Abbeel , Sergey Levine

We consider penalized extremum estimation of a high-dimensional, possibly nonlinear model that is sparse in the sense that most of its parameters are zero but some are not. We use the SCAD penalty function, which provides model selection…

计量经济学 · 经济学 2024-02-23 Joel L. Horowitz , Ahnaf Rafi

We apply the analytic conformal bootstrap method to study weakly coupled conformal gauge theories in four dimensions. We employ twist conformal blocks to find the most general form of the one-loop four-point correlation function of…

高能物理 - 理论 · 物理学 2018-04-04 Johan Henriksson , Tomasz Lukowski

ATARI is a suite of video games used by reinforcement learning (RL) researchers to test the effectiveness of the learning algorithm. Receiving only the raw pixels and the game score, the agent learns to develop sophisticated strategies,…

机器学习 · 计算机科学 2024-07-17 Le Zhang , Yong Gu , Xin Zhao , Yanshuo Zhang , Shu Zhao , Yifei Jin , Xinxin Wu

We introduce a sequence-conditioned critic for Soft Actor--Critic (SAC) that models trajectory context with a lightweight Transformer and trains on aggregated $N$-step targets. Unlike prior approaches that (i) score state--action pairs in…

机器学习 · 计算机科学 2025-09-30 Dong Tian , Onur Celik , Gerhard Neumann

We propose the relaxation bootstrap method for the numerical solution of multi-matrix models in the large $N$ limit, developing and improving the recent proposal of H.Lin. It gives rigorous inequalities on the single trace moments of the…

高能物理 - 理论 · 物理学 2022-06-22 Vladimir Kazakov , Zechuan Zheng

In this work we initiate a positive semi-definite numerical bootstrap program for multi-point correlators. Considering six-point functions of operators on a line we reformulate the crossing symmetry equation for a pair of comb-channel…

高能物理 - 理论 · 物理学 2024-06-19 António Antunes , Sebastian Harris , Apratim Kaviraj , Volker Schomerus

We propose WSAC (Weighted Safe Actor-Critic), a novel algorithm for Safe Offline Reinforcement Learning (RL) under functional approximation, which can robustly optimize policies to improve upon an arbitrary reference policy with limited…

机器学习 · 计算机科学 2024-11-01 Honghao Wei , Xiyue Peng , Arnob Ghosh , Xin Liu

We present Distributional Soft Actor-Critic (DSAC), a distributional reinforcement learning (RL) algorithm that combines the strengths of distributional information of accumulated rewards and entropy-driven exploration from Soft…

机器学习 · 计算机科学 2025-07-01 Xiaoteng Ma , Junyao Chen , Li Xia , Jun Yang , Qianchuan Zhao , Zhengyuan Zhou

Despite the empirical success of the actor-critic algorithm, its theoretical understanding lags behind. In a broader context, actor-critic can be viewed as an online alternating update algorithm for bilevel optimization, whose convergence…

机器学习 · 计算机科学 2019-07-16 Zhuoran Yang , Yongxin Chen , Mingyi Hong , Zhaoran Wang

Soft Actor-Critic is a state-of-the-art reinforcement learning algorithm for continuous action settings that is not applicable to discrete action settings. Many important settings involve discrete actions, however, and so here we derive an…

机器学习 · 计算机科学 2019-10-21 Petros Christodoulou