中文
相关论文

相关论文: Recurrent Structural Policy Gradient for Partially…

200 篇论文

In this work, we consider a novel inverse problem in mean-field games (MFG). We aim to recover the MFG model parameters that govern the underlying interactions among the population based on a limited set of noisy partial observations of the…

数值分析 · 数学 2022-04-12 Yat Tin Chow , Samy Wu Fung , Siting Liu , Levon Nurbekyan , Stanley Osher

This paper presents a Mean Field Game (MFG) model for maritime traffic flow, treating the navigation of ships between seaports as a large-scale stochastic control problem. The MFG framework enables the modeling of agents at a microscopic…

最优化与控制 · 数学 2025-12-02 Charles-Albert Lehalle , Giulia Livieri

Software reliability growth models (SRGM) enable failure data collected during testing. Specifically, nonhomogeneous Poisson process (NHPP) SRGM are the most commonly employed models. While software reliability growth models are important,…

软件工程 · 计算机科学 2024-02-01 Shadow Pritchard , Bhaskar Mitra , Vidhyashree Nagaraju

Optimizing with group sparsity is significant in enhancing model interpretability in machining learning applications, e.g., feature selection, compressed sensing and model compression. However, for large-scale stochastic training problems,…

最优化与控制 · 数学 2021-02-16 Tianyi Chen , Guanyi Wang , Tianyu Ding , Bo Ji , Sheng Yi , Zhihui Zhu

Mean-field games (MFG) were introduced to efficiently analyze approximate Nash equilibria in large population settings. In this work, we consider entropy-regularized mean-field games with a finite state-action space in a discrete time…

计算机科学与博弈论 · 计算机科学 2022-07-26 Yue Guan , Mi Zhou , Ali Pakniyat , Panagiotis Tsiotras

Finding Nash equilibria in two-player zero-sum imperfect-information games remains a central challenge in multi-agent reinforcement learning. Recent multi-round regularization methods offer a promising direction, yet existing approaches…

机器学习 · 计算机科学 2026-05-01 Eason Yu , Tzu Hao Liu , Clément L. Canonne , Yunke Wang , Chang Xu , Nguyen H. Tran , Stefano V. Albrecht

Medical time-series data are characterized by irregular sampling, high noise levels, missing values, and strong inter-feature dependencies. Recurrent neural networks (RNNs), particularly gated architectures such as Long Short-Term Memory…

机器学习 · 计算机科学 2026-03-03 Maitri Krishna Sai

Generative robot policies such as Flow Matching offer flexible, multi-modal policy learning but are sample-inefficient. Although object-centric policies improve sample efficiency, it does not resolve this limitation. In this work, we…

机器人学 · 计算机科学 2026-04-01 Jan Ole von Hartz , Lukas Schweizer , Joschka Boedecker , Abhinav Valada

We propose a novel hybrid stochastic policy gradient estimator by combining an unbiased policy gradient estimator, the REINFORCE estimator, with another biased one, an adapted SARAH estimator for policy optimization. The hybrid policy…

机器学习 · 计算机科学 2020-09-23 Nhan H. Pham , Lam M. Nguyen , Dzung T. Phan , Phuong Ha Nguyen , Marten van Dijk , Quoc Tran-Dinh

We introduce Group Policy Gradient (GPG), a family of critic-free policy-gradient estimators for general MDPs. Inspired by the success of GRPO's approach in Reinforcement Learning from Human Feedback (RLHF), GPG replaces a learned value…

机器学习 · 计算机科学 2025-10-07 Junhua Chen , Zixi Zhang , Hantao Zhong , Rika Antonova

In this work, we present an application of the probabilistic weak formulation of mean field games (MFG) for modeling liquidity pools in a constant product automated market maker (AMM) protocol in the context of decentralized finance. Our…

最优化与控制 · 数学 2026-04-14 Agustín Muñoz González , Juan I. Sequeira , Rafael Orive Illera

Reinforcement learning (RL) has achieved remarkable success in a wide range of control and decision-making tasks. However, RL agents often exhibit unstable or degraded performance when deployed in environments subject to unexpected external…

机器学习 · 计算机科学 2026-03-13 Taeho Lee , Donghwan Lee

Reinforcement learning algorithms for mean-field games offer a scalable framework for optimizing policies in large populations of interacting agents. Existing methods often depend on online interactions or access to system dynamics,…

机器学习 · 计算机科学 2024-10-24 Axel Brunnbauer , Julian Lemmel , Zahra Babaiee , Sophie Neubauer , Radu Grosu

This paper introduces a framework of Constrained Mean-Field Games (CMFGs), where each agent solves a constrained Markov decision process (CMDP). This formulation captures scenarios in which agents' strategies are subject to feasibility,…

最优化与控制 · 数学 2025-10-15 Anran Hu , Zijiu Lyu

Traditional model-based reinforcement learning approaches learn a model of the environment dynamics without explicitly considering how it will be used by the agent. In the presence of misspecified model classes, this can lead to poor…

机器学习 · 计算机科学 2020-10-20 Pierluca D'Oro , Alberto Maria Metelli , Andrea Tirinzoni , Matteo Papini , Marcello Restelli

We study discrete-time, finite-state mean-field games (MFGs) under model uncertainty, where agents face ambiguity about the state transition probabilities. Each agent maximizes its expected payoff against the worst-case transitions within…

最优化与控制 · 数学 2026-01-21 Zongxia Liang , Zhou Zhou , Yaqi Zhuang , Bin Zou

A growing line of work reframes preference-based fine-tuning of large language models game-theoretically: Nash Learning from Human Feedback (NLHF) recasts the problem as a zero-sum game over policies. However, optimization is over expected…

计算机科学与博弈论 · 计算机科学 2026-05-14 Max Horwitz , Jake Gonzales , Eric Mazumdar , Lillian J. Ratliff

Despite extreme sample inefficiency, on-policy reinforcement learning, aka policy gradients, has become a fundamental tool in decision-making problems. With the recent advances in GPU-driven simulation, the ability to collect large amounts…

机器学习 · 计算机科学 2024-07-30 Jayesh Singla , Ananye Agarwal , Deepak Pathak

In this article we consider finite Mean Field Games (MFGs), i.e. with finite time and finite states. We adopt the framework introduced in Gomes Mohr and Souza in 2010, and study two seemly unexplored subjects. In the first one, we analyze…

最优化与控制 · 数学 2018-05-16 Saeed Hadikhanloo , Francisco José Silva

The recent progress in multi-agent deep reinforcement learning(MADRL) makes it more practical in real-world tasks, but its relatively poor scalability and the partially observable constraints raise challenges to its performance and…

机器学习 · 计算机科学 2021-09-07 Zhenhui Ye , Xiaohong Jiang , Guanghua Song , Bowei Yang