中文
相关论文

相关论文: A Nash Equilibrium Framework For Training-Free Mul…

200 篇论文

In this paper, we propose a numerical methodology for finding the closed-loop Nash equilibrium of stochastic delay differential games through deep learning. These games are prevalent in finance and economics where multi-agent interaction…

最优化与控制 · 数学 2023-07-14 Robert Balkin , Hector D. Ceniceros , Ruimeng Hu

$S$ equilibrium synthesizes a century of game-theoretic modeling. $S$-beliefs determine choices as in the refinement literature and level-$k$, without anchoring on Nash equilibrium or imposing ad hoc belief formation. $S$-choices allow for…

理论经济学 · 经济学 2023-07-13 Jacob K Goeree , Bernardo Garcia-Pola

Outcome-rewarded Large Language Models (LLMs) have demonstrated remarkable success in mathematical problem-solving. However, this success often masks a critical issue: models frequently achieve correct answers through fundamentally unsound…

计算与语言 · 计算机科学 2025-06-25 Jiaxing Guo , Wenjie Yang , Shengzhong Zhang , Tongshan Xu , Lun Du , Da Zheng , Zengfeng Huang

Large language models (LLMs) have demonstrated impressive performance in various natural language processing tasks, yet their ability to perform multi-step logical reasoning remains an open challenge. Although Chain-of-Thought prompting has…

We develop a game-theoretic framework for predicting and steering the behavior of populations of large language models (LLMs) through Nash equilibrium (NE) analysis. To avoid the intractability of equilibrium computation in open-ended text…

人工智能 · 计算机科学 2026-02-09 Tonghan Wang , Yuqi Pan , Xinyi Yang , Yanchen Jiang , Milind Tambe , David C. Parkes

Finding Nash equilibria in two-player zero-sum imperfect-information games remains a central challenge in multi-agent reinforcement learning. Recent multi-round regularization methods offer a promising direction, yet existing approaches…

机器学习 · 计算机科学 2026-05-01 Eason Yu , Tzu Hao Liu , Clément L. Canonne , Yunke Wang , Chang Xu , Nguyen H. Tran , Stefano V. Albrecht

Algorithm design and analysis is a cornerstone of computer science, but it confronts a major challenge. Proving an algorithm's performance guarantee across all inputs has traditionally required extensive and often error-prone human effort.…

计算机科学与博弈论 · 计算机科学 2025-08-19 Hanyu Li , Dongchen Li , Xiaotie Deng

Reinforcement learning post-training has substantially improved the reasoning accuracy of vision-language models, yet the resulting policies remain poorly calibrated. Terminal correctness rewards provide no gradient that penalizes confident…

机器学习 · 计算机科学 2026-05-19 Peng Cui , Boyao Yang , Jun Zhu

Test-time scaling via solution sampling and aggregation has become a key paradigm for improving the reasoning performance of Large Language Models (LLMs). While reward model selection is commonly employed in this approach, it often fails to…

机器学习 · 计算机科学 2025-09-30 Zhicheng Yang , Zhijiang Guo , Yinya Huang , Yongxin Wang , Yiwei Wang , Xiaodan Liang , Jing Tang

This paper considers the problem of inverse reinforcement learning in zero-sum stochastic games when expert demonstrations are known to be not optimal. Compared to previous works that decouple agents in the game by assuming optimality in…

机器学习 · 统计学 2018-06-07 Xingyu Wang , Diego Klabjan

We introduce a set-valued solution concept, M equilibrium, to capture empirical regularities from over half a century of game-theory experiments. We show M equilibrium serves as a meta theory for various models that hitherto were considered…

理论经济学 · 经济学 2021-04-20 Jacob K. Goeree , Philippos Louis

LLM agents are known to deviate from Nash equilibria in strategic interactions, but nobody has looked inside the model to understand why, or asked whether the deviation can be reversed. We do both. Working with four open-source models…

计算机科学与博弈论 · 计算机科学 2026-05-05 Paraskevas V. Lekeas , Giorgos Stamatopoulos

Model-free learning for multi-agent stochastic games is an active area of research. Existing reinforcement learning algorithms, however, are often restricted to zero-sum games, and are applicable only in small state-action spaces or other…

机器学习 · 计算机科学 2022-10-25 Philippe Casgrain , Brian Ning , Sebastian Jaimungal

In this paper, we investigate a new problem called narrative action evaluation (NAE). NAE aims to generate professional commentary that evaluates the execution of an action. Unlike traditional tasks such as score-based action quality…

计算机视觉与模式识别 · 计算机科学 2024-04-29 Shiyi Zhang , Sule Bai , Guangyi Chen , Lei Chen , Jiwen Lu , Junle Wang , Yansong Tang

In Multi-task learning (MTL), a joint model is trained to simultaneously make predictions for several tasks. Joint training reduces computation costs and improves data efficiency; however, since the gradients of these different tasks may…

机器学习 · 计算机科学 2022-07-11 Aviv Navon , Aviv Shamsian , Idan Achituve , Haggai Maron , Kenji Kawaguchi , Gal Chechik , Ethan Fetaya

Verifying multi-step reasoning in large language models is difficult due to imprecise error localization and high token costs. Existing methods either assess entire reasoning chains, suffering attention dilution, or rely on expensive…

人工智能 · 计算机科学 2025-10-06 Yulong Zhang , Li Wang , Wei Du , Peilin Li , Yuqin Dai Zhiyuan Zhao , Lingyong Fang , Ziniu Liu , Ru Zhang , Huijia Zhu , Gongshen Liu

This paper addresses the problem of learning a Nash equilibrium in $\gamma$-discounted multiplayer general-sum Markov Games (MG). A key component of this model is the possibility for the players to either collaborate or team apart to…

计算机科学与博弈论 · 计算机科学 2017-03-07 Julien Pérolat , Florian Strub , Bilal Piot , Olivier Pietquin

We introduce a novel class of Nash equilibrium seeking dynamics for non-cooperative games with a finite number of players, where the convergence to the Nash equilibrium is bounded by a KL function with a settling time that can be upper…

最优化与控制 · 数学 2020-12-25 Jorge I. Poveda , Miroslav Krstic , Tamer Basar

The design of Nash equilibrium seeking strategies for games in which the involved players are of second-order integrator-type dynamics is investigated in this paper. Noticing that velocity signals are usually noisy or not available for…

最优化与控制 · 数学 2020-06-18 Maojiao Ye , Jizhao Yin , Le Yin

Reasoning is an essential capacity for large language models (LLMs) to address complex tasks, where the identification of process errors is vital for improving this ability. Recently, process-level reward models (PRMs) were proposed to…

人工智能 · 计算机科学 2025-03-18 Zhaopan Xu , Pengfei Zhou , Jiaxin Ai , Wangbo Zhao , Kai Wang , Xiaojiang Peng , Wenqi Shao , Hongxun Yao , Kaipeng Zhang