中文
相关论文

相关论文: Diverse Exploration via Conjugate Policies for Pol…

200 篇论文

How to incentivize self-interested agents to explore when they prefer to exploit? Consider a population of self-interested agents that make decisions under uncertainty. They "explore" to acquire new information and "exploit" this…

计算机科学与博弈论 · 计算机科学 2024-10-23 Aleksandrs Slivkins

Intrinsically motivated goal exploration processes enable agents to autonomously sample goals to explore efficiently complex environments with high-dimensional continuous actions. They have been applied successfully to real world robots to…

机器学习 · 计算机科学 2018-11-06 Adrien Laversanne-Finot , Alexandre Péré , Pierre-Yves Oudeyer

The Centralized Training with Decentralized Execution (CTDE) paradigm is widely used in cooperative multi-agent reinforcement learning. However, conventional methods based on CTDE can suffer from value underestimation and converge to…

多智能体系统 · 计算机科学 2026-05-05 Ruoning Zhang , Siying Wang , Wenyu Chen , Yang Zhou , Zhitong Zhao , Zixuan Zhang , Ruijie Zhang , Stefano V. Albrecht

Policy-gradient methods such as Proximal Policy Optimization (PPO) are typically updated along a single stochastic gradient direction, leaving the rich local structure of the parameter space unexplored. Previous work has shown that the…

机器学习 · 计算机科学 2025-10-01 Xinyu Zhang , Aishik Deb , Klaus Mueller

This perspective posits that gene-environment interplay (GxE) studies should be developed both theoretically and empirically to be of relevance to policy makers. On the theoretical front, this development is essential because the current…

其他定量生物学 · 定量生物学 2025-07-17 Dilnoza Muslimova , Niels Rietveld

The optimization of expensive black-box simulators arises in a myriad of modern scientific and engineering applications. Bayesian optimization provides an appealing solution, by leveraging a fitted surrogate model to guide the selection of…

Policy gradient methods are powerful reinforcement learning algorithms and have been demonstrated to solve many complex tasks. However, these methods are also data-inefficient, afflicted with high variance gradient estimates, and frequently…

机器学习 · 计算机科学 2019-05-15 Andreas Doerr , Michael Volpp , Marc Toussaint , Sebastian Trimpe , Christian Daniel

Crossover is a powerful mechanism for generating new solutions from a given population of solutions. Crossover comes with a discrepancy in itself: on the one hand, crossover usually works best if there is enough diversity in the population;…

神经与进化计算 · 计算机科学 2025-07-03 Johannes Lengler , Tom Offermann

A numerical method for coupled 3D-1D problems with discontinuous solutions at the interfaces is derived and discussed. This extends a previous work on the subject where only continuous solutions were considered. Thanks to properly defined…

数值分析 · 数学 2022-03-04 Stefano Berrone , Denise Grappein , Stefano Scialò

The ability to differentiate through optimization problems has unlocked numerous applications, from optimization-based layers in machine learning models to complex design problems formulated as bilevel programs. It has been shown that…

最优化与控制 · 数学 2024-03-05 Lucas Fuentes Valenzuela , Robin Brown , Marco Pavone

Autonomous 3D environment exploration is a fundamental task for various applications such as navigation. The goal of exploration is to investigate a new environment and build its occupancy map efficiently. In this paper, we propose a new…

人工智能 · 计算机科学 2021-11-03 Liu Juncheng , McCane Brendan , Mills Steven

This paper studies privacy in the context of complex decision support queries composed of multiple conditions on different aggregate statistics combined using disjunction and conjunction operators. Utility requirements for such queries…

数据库 · 计算机科学 2024-06-25 Nada Lahjouji , Sameera Ghayyur , Xi He , Sharad Mehrotra

We propose a method to accurately and efficiently identify the constitutive behavior of complex materials through full-field observations. We formulate the problem of inferring constitutive relations from experiments as an indirect inverse…

材料科学 · 物理学 2024-12-05 Andrew Akerson , Aakila Rajan , Kaushik Bhattacharya

Direct policy optimization in reinforcement learning is usually solved with policy-gradient algorithms, which optimize policy parameters via stochastic gradient ascent. This paper provides a new theoretical interpretation and justification…

机器学习 · 计算机科学 2023-10-24 Adrien Bolland , Gilles Louppe , Damien Ernst

Differentiable conjugacies link dynamical systems that share properties such as the stability multipliers of corresponding orbits. It provides a stronger classification than topological conjugacy, which only requires qualitative similarity.…

动力系统 · 数学 2023-03-02 P. A. Glendinning , D. J. W. Simpson

I development a Conjugate Gradient Method for solving a partial differential system with multiply controls. Some numerical results are depicted. Also, I present an explication of why the control over a partial differential equations system…

最优化与控制 · 数学 2014-11-25 Carlos Barrón-Romero

We consider a team of reinforcement learning agents that concurrently learn to operate in a common environment. We identify three properties - adaptivity, commitment, and diversity - which are necessary for efficient coordinated exploration…

人工智能 · 计算机科学 2018-12-18 Maria Dimakopoulou , Benjamin Van Roy

This paper addresses unconstrained multiobjective optimization problems where two or more continuously differentiable functions have to be minimized. We delve into the conjugate gradient methods proposed by Lucambio P\'{e}rez and Prudente…

最优化与控制 · 数学 2024-10-15 Wang Chen , Yong Zhao , Liping Tang , Xinmin Yang

Complex single-objective bounded problems are often difficult to solve. In evolutionary computation methods, since the proposal of differential evolution algorithm in 1997, it has been widely studied and developed due to its simplicity and…

神经与进化计算 · 计算机科学 2024-04-26 Sichen Tao , Ruihan Zhao , Kaiyu Wang , Shangce Gao

This paper considers policy search in continuous state-action reinforcement learning problems. Typically, one computes search directions using a classic expression for the policy gradient called the Policy Gradient Theorem, which decomposes…

机器学习 · 计算机科学 2020-04-13 Sujay Bhatt , Alec Koppel , Vikram Krishnamurthy