English
Related papers

Related papers: Reducing the Incentive to Tank: The Ex Post Gold P…

200 papers

Stemming on the idea that a key objective in reinforcement learning is to invert a target distribution of effects, end-effect drives are proposed as an effective way to implement goal-directed motor learning, in the absence of an explicit…

Artificial Intelligence · Computer Science 2020-10-06 Emmanuel Daucé

We propose a restricted win probability estimand for comparing treatments in a randomized trial with a time-to-event outcome. We also propose Bayesian estimators for this summary measure as well as the unrestricted win probability. Bayesian…

Methodology · Statistics 2024-11-06 Michelle Leeberg , Xianghua Luo , Thomas A. Murray

We consider the problem of incentivising desirable behaviours in multi-agent systems by way of taxation schemes. Our study employs the concurrent games model: in this model, each agent is primarily motivated to seek the satisfaction of a…

Computer Science and Game Theory · Computer Science 2023-07-12 David Hyland , Julian Gutierrez , Michael Wooldridge

The paper discusses the strategy-proofness of sports tournaments with multiple group stages, where the results of matches already played in the previous round against teams in the same group are carried over. These tournaments, widely used…

Physics and Society · Physics 2023-02-24 László Csató

This paper proposes a new reinforcement learning with hyperbolic discounting. Combining a new temporal difference error with the hyperbolic discounting in recursive manner and reward-punishment framework, a new scheme to learn the optimal…

Machine Learning · Computer Science 2021-06-04 Taisuke Kobayashi

We study the effects of randomness on competitions based on an elementary random process in which there is a finite probability that a weaker team upsets a stronger team. We apply this model to sports leagues and sports tournaments, and…

Physics and Society · Physics 2013-04-02 E. Ben-Naim , N. W. Hengartner , S. Redner , F. Vazquez

Current reinforcement learning objectives for large-model reasoning primarily focus on maximizing expected rewards. This paradigm can lead to overfitting to dominant reward signals, while neglecting alternative yet valid reasoning…

Machine Learning · Computer Science 2026-02-24 Wendi Li , Sharon Li

Reinforcement learning (RL) is the dominant paradigm for post-training large language models. However, in the online, on-policy setting, rollout generation dominates the computational cost of training. Group-based policy optimization…

Machine Learning · Computer Science 2026-05-27 Woojeong Kim , Ziyi Yang , Jing Nathan Yan , Jialu Liu

Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regularization with respect to the base policy. This practice is suboptimal for maximizing user's utility…

Machine Learning · Computer Science 2026-02-04 Haichuan Wang , Tao Lin , Lingkai Kong , Ce Li , Hezi Jiang , Milind Tambe

A growing body of evidence has shown that incorporating behavioral economics principles into the design of financial incentive programs helps improve their cost-effectiveness, promote individuals' short-term engagement, and increase…

Social and Information Networks · Computer Science 2020-10-28 Palakorn Achananuparp , Ee-Peng Lim , Vibhanshu Abhishek , Tianjiao Yun

Fair re-ranking aims to redistribute ranking slots among items more equitably to ensure responsibility and ethics. The exploration of redistribution problems has a long history in economics, offering valuable insights for conceptualizing…

Information Retrieval · Computer Science 2024-04-30 Chen Xu , Xiaopeng Ye , Wenjie Wang , Liang Pang , Jun Xu , Tat-Seng Chua

Aligning large language models (LLMs) with human preferences has become a critical step in their development. Recent research has increasingly focused on test-time alignment, where additional compute is allocated during inference to enhance…

Computation and Language · Computer Science 2025-09-24 Bolian Li , Yanran Wu , Xinyu Luo , Ruqi Zhang

Discounted-sum games provide a formal model for the study of reinforcement learning, where the agent is enticed to get rewards early since later rewards are discounted. When the agent interacts with the environment, she may regret her…

Computer Science and Game Theory · Computer Science 2018-11-20 Michaël Cadilhac , Guillermo A. Pérez , Marie van den Bogaard

The outputs of win probability models are often used to evaluate player actions. However, in some sports, such as the popular esport Counter-Strike, there exist important team-level decisions. For example, at the beginning of each round in…

Computer Science and Game Theory · Computer Science 2021-09-28 Peter Xenopoulos , Bruno Coelho , Claudio Silva

We present the first reinforcement-learning model to self-improve its reward-modulated training implemented through a continuously improving "intuition" neural network. An agent was trained how to play the arcade video game Pong with two…

Artificial Intelligence · Computer Science 2016-09-26 Matt Oberdorfer , Matt Abuzalaf

If the final position of a team is already secured independently of the outcomes of the remaining games in a round-robin tournament, it might play with little enthusiasm. This is detrimental to attendance and can inspire collusion and…

Physics and Society · Physics 2023-02-03 László Csató

We introduce penalty-function-based admission control policies to approximately maximize the expected reward rate in a loss network. These control policies are easy to implement and perform well both in the transient period as well as in…

Probability · Mathematics 2007-05-23 Garud Iyengar , Karl Sigman

In many competitive settings, from education to politics, rules do not reward effort evenly, and thresholds (e.g., grade cutoffs or electoral majorities) make some moments disproportionately important. Success thus depends on efficiently…

Computer Science and Game Theory · Computer Science 2026-01-23 Masatsugu Yoshizawa , Yuta Kawamoto , Daisuke Takeshita

This paper examines strategic effort and positioning choices resulting in bandwagon effects under externalities in finite multi-stage games using causal evidence from triathlon (Reichel, 2025). Focusing on open-water swim drafting where…

General Economics · Economics 2025-05-13 Felix Reichel

Reinforcement learning agents for portfolio management are typically trained and deployed as static policies, with no mechanism for using price forecasts at inference time. We propose $\text{FPILOT}$ (**Fin**ancial **P**lugin…

Machine Learning · Computer Science 2026-05-14 Eun Go , Rohan Deb , Arindam Banerjee