中文
相关论文

相关论文: Discounting in Strategy Logic

200 篇论文

Humans exhibit time-inconsistent behavior, in which planned actions diverge from executed actions. Understanding time inconsistency and designing appropriate interventions is a key research challenge in computer science and behavioral…

计算机科学与博弈论 · 计算机科学 2025-09-18 Yasunori Akagi , Takeshi Kurashima

Fine-tuning large language models (LLMs) to aggregate multiple preferences has attracted considerable research attention. With aggregation algorithms advancing, a potential economic scenario arises where fine-tuning services are provided to…

计算机科学与博弈论 · 计算机科学 2026-02-11 Haoran Sun , Yurong Chen , Siwei Wang , Xu Chu , Wei Chen , Xiaotie Deng

This paper is an original attempt to understand the foundations of economic reasoning. It endeavors to rigorously define the relationship between subjective interpretations and objective valuations of such interpretations in the context of…

计算机科学中的逻辑 · 计算机科学 2024-05-20 Daniel Lu

Future reward estimation is a core component of reinforcement learning agents; i.e., Q-value and state-value functions, predicting an agent's sum of future rewards. Their scalar output, however, obfuscates when or what individual future…

人工智能 · 计算机科学 2024-08-16 Mark Towers , Yali Du , Christopher Freeman , Timothy J. Norman

We study deterministic, discrete linear time-invariant systems with infinite-horizon discounted quadratic cost. It is well-known that standard stabilizability and detectability properties are not enough in general to conclude stability…

最优化与控制 · 数学 2025-09-04 Jonathan de Brusse , Jamal Daafouz , Mathieu Granzotto , Romain Postoyan , Dragan Nesic

The optimal objective is a fundamental aspect of reinforcement learning (RL), as it determines how policies are evaluated and optimized. While total return maximization is the ideal objective in RL, discounted return maximization is the…

机器学习 · 计算机科学 2025-03-19 Shuyu Yin , Fei Wen , Peilin Liu , Tao Luo

We consider decision-making and game scenarios in which an agent is limited by his/her computational ability to foresee all the available moves towards the future - that is, we study scenarios with short sight. We focus on how short sight…

计算机科学中的逻辑 · 计算机科学 2016-06-27 Chanjuan Liu

Strategy logic (SL) is a powerful temporal logic that enables strategic reasoning in multi-agent systems. SL supports explicit (first-order) quantification over strategies and provides a logical framework to express many important…

多智能体系统 · 计算机科学 2024-03-21 Raven Beutner , Bernd Finkbeiner

We explore the effect of discounting and experimentation in a simple model of interacting adaptive agents. Agents belong to either of two types and each has to decide whether to participate a game or not, the game being profitable when…

物理与社会 · 物理学 2009-11-13 Damien Challet , Andrea De Martino , Matteo Marsili

Data analysis and performance evaluation of simulation deduction plays a pivotal role in modern warfare, which enables military personnel to gain invaluable insights into the potential effectiveness of different strategies, tactics, and…

计算与语言 · 计算机科学 2025-11-17 Shansi Zhang , Min Li

People often face trade-offs between costs and benefits occurring at various points in time. The predominant discounting approach is to use the exponential form. Central to this approach is the discount rate, a unique parameter that…

理论经济学 · 经济学 2024-08-13 Bach Dong-Xuan , Philippe Bich

Many popular policy gradient methods for reinforcement learning follow a biased approximation of the policy gradient known as the discounted approximation. While it has been shown that the discounted approximation of the policy gradient is…

机器学习 · 计算机科学 2023-01-10 Chris Nota

Time-inconsistent preferences, where agents favor smaller-sooner over larger-later rewards, are a key feature of human and animal decision-making. Quasi-Hyperbolic (QH) discounting provides a simple yet powerful model for this behavior, but…

机器学习 · 计算机科学 2025-09-09 S. R. Eshwar

The policy improvement bound on the difference of the discounted returns plays a crucial role in the theoretical justification of the trust-region policy optimization (TRPO) algorithm. The existing bound leads to a degenerate bound when the…

机器学习 · 计算机科学 2021-07-20 J. G. Dai , Mark Gluzman

We investigate the discounting mismatch in actor-critic algorithm implementations from a representation learning perspective. Theoretically, actor-critic algorithms usually have discounting for both actor and critic, i.e., there is a…

机器学习 · 计算机科学 2022-01-27 Shangtong Zhang , Romain Laroche , Harm van Seijen , Shimon Whiteson , Remi Tachet des Combes

Empirical research often cites observed choice responses to variation that shifts expected discounted future utilities, but not current utilities, as an intuitive source of information on time preferences. We study the identification of…

计量经济学 · 经济学 2020-05-28 Jaap H. Abbring , Øystein Daljord

Processes (MDPs) often require frequent decision making, that is, taking an action every microsecond, second, or minute. Infinite horizon discount reward formulation is still relevant for a large portion of these applications, because…

最优化与控制 · 数学 2014-12-17 Yin-Lam Chow , Junjie Qin

Our aim is to design mechanisms that motivate all agents to reveal their predictions truthfully and promptly. For myopic agents, proper scoring rules induce truthfulness. However, as has been described in the literature, when agents take…

计算机科学与博弈论 · 计算机科学 2019-12-05 Amir Ban

This paper proposes a new reinforcement learning with hyperbolic discounting. Combining a new temporal difference error with the hyperbolic discounting in recursive manner and reward-punishment framework, a new scheme to learn the optimal…

机器学习 · 计算机科学 2021-06-04 Taisuke Kobayashi

In decision support systems, it is essential to get a candidate solution fast, even if it means resorting to an approximation. This constraint introduces a scalability requirement with regard to the kind of heuristics which can be used in…

多智能体系统 · 计算机科学 2014-05-22 D. Krzywicki , Ł. Faber , A. Byrski , M. Kisiel-Dorohinicki