中文
相关论文

相关论文: Last-iterate convergence of modified predictive me…

200 篇论文

Several widely-used first-order saddle-point optimization methods yield an identical continuous-time ordinary differential equation (ODE) that is identical to that of the Gradient Descent Ascent (GDA) method when derived naively. However,…

最优化与控制 · 数学 2023-08-01 Tatjana Chavdarova , Michael I. Jordan , Manolis Zampetakis

Min-max formulations have attracted great attention in the ML community due to the rise of deep generative models and adversarial methods, while understanding the dynamics of gradient algorithms for solving such formulations has remained a…

机器学习 · 计算机科学 2020-03-05 Guojun Zhang , Yaoliang Yu

Self-play methods have demonstrated remarkable success in enhancing model capabilities across various domains. In the context of Reinforcement Learning from Human Feedback (RLHF), self-play not only boosts Large Language Model (LLM)…

计算与语言 · 计算机科学 2025-04-22 Mingzhi Wang , Chengdong Ma , Qizhi Chen , Linjian Meng , Yang Han , Jiancong Xiao , Zhaowei Zhang , Jing Huo , Weijie J. Su , Yaodong Yang

The multireference alignment problem consists of estimating a signal from multiple noisy shifted observations. Inspired by existing Unique-Games approximation algorithms, we provide a semidefinite program (SDP) based relaxation which…

数据结构与算法 · 计算机科学 2013-08-27 Afonso S. Bandeira , Moses Charikar , Amit Singer , Andy Zhu

Last-iterate behaviors of learning algorithms in repeated two-player zero-sum games have been extensively studied due to their wide applications in machine learning and related tasks. Typical algorithms that exhibit the last-iterate…

机器学习 · 计算机科学 2024-06-18 Yi Feng , Ping Li , Ioannis Panageas , Xiao Wang

We study last-iterate convergence properties of algorithms for solving two-player zero-sum games based on Regret Matching$^+$ (RM$^+$). Despite their widespread use for solving real games, virtually nothing is known about their last-iterate…

计算机科学与博弈论 · 计算机科学 2025-03-05 Yang Cai , Gabriele Farina , Julien Grand-Clément , Christian Kroer , Chung-Wei Lee , Haipeng Luo , Weiqiang Zheng

Our work focuses on extra gradient learning algorithms for finding Nash equilibria in bilinear zero-sum games. The proposed method, which can be formally considered as a variant of Optimistic Mirror Descent…

计算机科学与博弈论 · 计算机科学 2022-03-09 Michail Fasoulakis , Evangelos Markakis , Yannis Pantazis , Constantinos Varsos

While classic work in convex-concave min-max optimization relies on average-iterate convergence results, the emergence of nonconvex applications such as training Generative Adversarial Networks has led to renewed interest in last-iterate…

最优化与控制 · 数学 2019-10-29 Jacob Abernethy , Kevin A. Lai , Andre Wibisono

Last-iterate convergence has received extensive study in two player zero-sum games starting from bilinear, convex-concave up to settings that satisfy the MVI condition. Typical methods that exhibit last-iterate convergence for the…

计算机科学与博弈论 · 计算机科学 2023-10-05 Yi Feng , Hu Fu , Qun Hu , Ping Li , Ioannis Panageas , Bo Peng , Xiao Wang

The success of adversarial formulations in machine learning has brought renewed motivation for smooth games. In this work, we focus on the class of stochastic Hamiltonian methods and provide the first convergence guarantees for certain…

We introduce a frequency-domain framework for convergence analysis of hyperparameters in game optimization, leveraging High-Resolution Differential Equations (HRDEs) and Laplace transforms. Focusing on the Lookahead algorithm--characterized…

最优化与控制 · 数学 2025-06-17 Aniket Sanyal , Tatjana Chavdarova

We study policy optimization for infinite-horizon, discounted constrained Markov decision processes (CMDPs). While existing theoretical guarantees typically hold for the mixture policy, deploying such a policy is computationally and memory…

机器学习 · 计算机科学 2026-05-13 Michael Lu , Max Qiushi Lin , Mo Chen , Sharan Vaswani

This paper presents a class of evolutive Mean Field Games with multiple solutions for all time horizons T and convex but non-smooth Hamiltonian H, as well as for smooth H and T large enough. The phenomenon is analyzed in both the PDE and…

偏微分方程分析 · 数学 2018-02-12 Martino Bardi , Markus Fischer

We study convergence rates of the generalized conditional gradient (GCG) method applied to fully discretized Mean Field Games (MFG) systems. While explicit convergence rates of the GCG method have been established at the continuous PDE…

数值分析 · 数学 2026-02-13 Haruka Nakamura , Norikazu Saito

This paper proposes a payoff perturbation technique for the Mirror Descent (MD) algorithm in games where the gradient of the payoff functions is monotone in the strategy profile space, potentially containing additive noise. The optimistic…

计算机科学与博弈论 · 计算机科学 2024-06-25 Kenshi Abe , Kaito Ariu , Mitsuki Sakamoto , Atsushi Iwasaki

We introduce a near-linear complexity (geometric and meshless/algebraic) multigrid/multiresolution method for PDEs with rough ($L^\infty$) coefficients with rigorous a-priori accuracy and performance estimates. The method is discovered…

数值分析 · 数学 2017-02-13 Houman Owhadi

Mirror play (MP) is a well-accepted primal-dual multi-agent learning algorithm where all agents simultaneously implement mirror descent in a distributed fashion. The advantage of MP over vanilla gradient play lies in its usage of mirror…

计算机科学与博弈论 · 计算机科学 2024-03-26 Yunian Pan , Tao Li , Quanyan Zhu

This paper considers the problem of finding a solution to the finite horizon constrained Markov decision processes (CMDP) where the objective as well as constraints are sum of additive and multiplicative utilities. Towards solving this, we…

最优化与控制 · 数学 2023-03-16 Uday Kumar M , Sanjay P Bhat , Veeraruna Kavitha , Nandyala Hemachandra

An iterative finite difference scheme for mean field games (MFGs) is proposed. The target MFGs are derived from control problems for multidimensional systems with advection terms. For such MFGs, linearization using the Cole-Hopf…

最优化与控制 · 数学 2023-04-26 Daisuke Inoue , Yuji Ito , Takahito Kashiwabara , Norikazu Saito , Hiroaki Yoshida

This paper proposes an asymmetric perturbation technique for solving bilinear saddle-point optimization problems, commonly arising in minimax problems, game theory, and constrained optimization. Perturbing payoffs or values is known to be…

最优化与控制 · 数学 2026-02-16 Kenshi Abe , Mitsuki Sakamoto , Kaito Ariu , Atsushi Iwasaki
‹ 上一页 1 2 3 10 下一页 ›