Related papers: Reducing the Incentive to Tank: The Ex Post Gold P…
This paper introduces a novel method of adding intrinsic bonuses to task-oriented reward function in order to efficiently facilitate reinforcement learning search. While various bonuses have been designed to date, they are analogous to the…
Prompting methods recently achieve impressive success in few-shot learning. These methods modify input samples with prompt sentence pieces, and decode label tokens to map samples to corresponding labels. However, such a paradigm is very…
Recent work has shown that standard training via empirical risk minimization (ERM) can produce models that achieve high accuracy on average but low accuracy on underrepresented groups due to the prevalence of spurious features. A…
In this paper, we investigate the effectiveness of the home team bunting in extra innings of Major League Baseball games when the game is tied in the bottom of the inning. Using methods rooted in causal inference, we show that teams choose…
Lottery Ticket Hypothesis (LTH) raises keen attention to identifying sparse trainable subnetworks, or winning tickets, which can be trained in isolation to achieve similar or even better performance compared to the full models. Despite many…
In the game-theoretic model war of attrition, players are subject to an explicit cost proportional to the duration of contests. We construct a model where the time cost is not explicitly given, but instead depends implicitly on the…
We present here a simple mathematical model that provides a successful strategy, quantitatively, to ending the continued championship futility experienced by Canadian Hockey Teams. Competitive Intransitivity is used here as a simple…
Pruning is a well-established technique for removing unnecessary structure from neural networks after training to improve the performance of inference. Several recent results have explored the possibility of pruning at initialization time…
We study the design of effort-maximizing grading schemes between agents with private abilities. Assuming agents derive value from the information their grade reveals about their ability, we find that more informative grading schemes induce…
Efficiently allocating treatments with a budget constraint constitutes an important challenge across various domains. In marketing, for example, the use of promotions to target potential customers and boost conversions is limited by the…
In this paper, a new continuous scoring system for soccer is proposed, based on the proportion of time that a team is winning, losing or tied. Several simulations are made applying this technique to complete seasons of different leagues. As…
We study the optimal allocation of prizes in rank-order tournaments with loss averse agents. Prize sharing becomes increasingly optimal with loss aversion because more equitable prizes reduce the marginal psychological cost of anticipated…
The allocation of scarce donor organs constitutes one of the most consequential algorithmic challenges in healthcare. While the field is rapidly transitioning from rigid, rule-based systems to machine learning and data-driven optimization,…
Language model (LM) post-training (or alignment) involves maximizing a reward function that is derived from preference annotations. Direct Preference Optimization (DPO) is a popular offline alignment method that trains a policy directly on…
We propose and study an evolutionary minority game (EMG) in which the agents are allowed to choose among three possible options. Unlike the original EMG where the agents either win or lose one unit of wealth, the present model assigns one…
Ensembling is a popular method used to improve performance as a last resort. However, ensembling multiple models finetuned from a single pretrained model has been not very effective; this could be due to the lack of diversity among ensemble…
Tournament-based compensation schemes with forced distributions represent a widely adopted class of relative performance evaluation mechanisms in technology and corporate environments. These systems mandate within-team ranking and fixed…
Injecting new reasoning knowledge into Large Language Models (LLMs) via post-training often induces catastrophic forgetting. Recent studies emphasize the importance of on-policy data but suggest that KL-divergence fails to mitigate…
While foundation models have been exploited for various expert tasks through fine-tuning, any foundation model will become outdated due to its old knowledge or limited capability. Thus the underlying foundation model should be eventually…
The Team Orienteering Problem (TOP) is an NP-hard routing problem in which a fleet of identical vehicles aims at collecting rewards (prizes) available at given locations, while satisfying restrictions on the travel times. In TOP, each…