中文
相关论文

相关论文: A Nash Equilibrium Framework For Training-Free Mul…

200 篇论文

Progress in machine learning is measured by careful evaluation on problems of outstanding common interest. However, the proliferation of benchmark suites and environments, adversarial attacks, and other complications has diluted the basic…

机器学习 · 计算机科学 2018-11-01 David Balduzzi , Karl Tuyls , Julien Perolat , Thore Graepel

Large language models are increasingly used to support high-stakes decisions, potentially influencing who is granted bail or receives a loan. Naive chain-of-thought sampling can improve average decision accuracy, but has also been shown to…

机器学习 · 计算机科学 2025-07-16 Zara Hall , Melanie Subbiah , Thomas P Zollo , Kathleen McKeown , Richard Zemel

We investigate the set of Nash equilibrium payoffs for two person differential games. The main result of the paper is the characterization of the set of Nash equilibrium payoffs in the terms of nonsmooth analysis. Also we obtain the…

最优化与控制 · 数学 2015-03-17 Yurii Averboukh

The evaluation and post-training of large language models (LLMs) rely on supervision, but strong supervision for difficult tasks is often unavailable, especially when evaluating frontier models. In such cases, models are demonstrated to…

机器学习 · 计算机科学 2026-01-29 Tianyi Alex Qiu , Micah Carroll , Cameron Allen

This paper is devoted to a high-dimensional mixed leadership stochastic differential game on a finite horizon in feedback information mode, where the control variables enter into the diffusion term of state equation. A verification theorem…

最优化与控制 · 数学 2022-11-28 Qi Huang , Jingtao Shi

We design and analyze attention games that incentivize validators to check computation results. We show that no pure strategy Nash equilibrium of the game without outside parties exists by a simple argument. We then proceed to calculate the…

计算机科学与博弈论 · 计算机科学 2023-08-08 Akaki Mamageishvili , Edward W. Felten

We explore the use of policy approximations to reduce the computational cost of learning Nash equilibria in zero-sum stochastic games. We propose a new Q-learning type algorithm that uses a sequence of entropy-regularized soft policies to…

机器学习 · 计算机科学 2021-06-29 Yue Guan , Qifan Zhang , Panagiotis Tsiotras

Nash equilibrium is a fundamental solution concept in extensive-form games, while its efficient computation is still far from straightforward. This paper considers finite $n$-player extensive-form games with perfect recall under the…

计算机科学与博弈论 · 计算机科学 2026-04-15 Yuqing Hou

Reinforcement Learning (RL) has emerged as a pivotal mechanism for enhancing the complex reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevailing paradigms typically rely on solitary rollout strategies where…

计算与语言 · 计算机科学 2026-02-05 Lingzhuang Sun , Ruitong Liu , Yuxia Zhu , Xiaohan Xu , Jingxuan Wei , Xiangxiang Zhang , Bihui Yu , Wentao Zhang

When applied to question answering and other text generation tasks, language models (LMs) may be queried generatively (by sampling answers from their output distribution) or discriminatively (by using them to score or rank a set of…

计算机科学与博弈论 · 计算机科学 2023-10-16 Athul Paul Jacob , Yikang Shen , Gabriele Farina , Jacob Andreas

Equilibria of realistic multiplayer games constitute a key solution concept both in practical applications, such as online advertising auctions and electricity markets, and in analytical frameworks used to study strategic voting in…

计算机科学与博弈论 · 计算机科学 2025-11-18 Jakub Černý , Shuvomoy Das Gupta , Christian Kroer

The Nash equilibrium problem is a widely used tool to model non-cooperative games. Many solution methods have been proposed in the literature to compute solutions of Nash equilibrium problems with continuous strategy sets, but, besides some…

最优化与控制 · 数学 2015-12-03 Simone Sagratella

We consider a general-sum N-player linear-quadratic game with stochastic dynamics over a finite horizon and prove the global convergence of the natural policy gradient method to the Nash equilibrium. In order to prove the convergence of the…

最优化与控制 · 数学 2022-08-16 Ben Hambly , Renyuan Xu , Huining Yang

In this paper, we study closed-loop strong equilibrium strategies for the time-inconsistent control problem with higher-order moments formulated by [Wang et al. SIAM J. Control. Optim., 63 (2025), 1560--1589]. Since time-inconsistency makes…

最优化与控制 · 数学 2025-08-15 Yike Wang

In this work, we study stochastic non-cooperative games, where only noisy black-box function evaluations are available to estimate the cost function for each player. Since each player's cost function depends on both its own decision…

计算机科学与博弈论 · 计算机科学 2025-11-18 Haidong Li , Anzhi Sheng , Yijie Peng , Long Wang

We consider shared workspace scenarios with humans and robots acting to achieve independent goals, termed as parallel play. We model these as general-sum games and construct a framework that utilizes the Nash equilibrium solution concept to…

人工智能 · 计算机科学 2020-06-11 Shray Bansal , Jin Xu , Ayanna Howard , Charles Isbell

Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating queries without considering user experiences, inadequately…

Large language model (LLM)-based reasoning systems have recently achieved gold medal-level performance in the IMO 2025 competition, writing mathematical proofs where, to receive full credit, each step must be not only correct but also…

人工智能 · 计算机科学 2025-10-16 Shrey Pandit , Austin Xu , Xuan-Phi Nguyen , Yifei Ming , Caiming Xiong , Shafiq Joty

Recent progress in large language models (LLM) found chain-of-thought prompting strategies to improve the reasoning ability of LLMs by encouraging problem solving through multiple steps. Therefore, subsequent research aimed to integrate the…

计算与语言 · 计算机科学 2025-02-21 Ting-Ruen Wei , Haowei Liu , Xuyang Wu , Yi Fang

We consider payoff-based learning of a generalized Nash equilibrium (GNE) in multi-agent systems. Our focus is on games with jointly convex constraints of a linear structure and strongly monotone pseudo-gradients. We present a convergent…

最优化与控制 · 数学 2025-07-18 Tatiana Tatarenko , Maryam Kamgarpour