English
Related papers

Related papers: Reducing the Incentive to Tank: The Ex Post Gold P…

200 papers

Speculative decoding has emerged as a pivotal technique to accelerate LLM inference by employing a lightweight draft model to generate candidate tokens that are subsequently verified by the target model in parallel. However, while this…

Computation and Language · Computer Science 2026-02-26 Yuetao Chen , Xuliang Wang , Xinzhou Zheng , Ming Li , Peng Wang , Hong Xu

We introduce a new class of balanced allocation processes which bias towards underloaded bins (those with load below the mean load) either by skewing the probability by which a bin is chosen for an allocation (probability bias), or…

Probability · Mathematics 2024-01-12 Dimitrios Los , Thomas Sauerwald , John Sylvester

We evaluate the tendency for different voting methods to promote political compromise and reduce tensions in a society by using computer simulations to determine which voters candidates are incentivized to appeal to. We find that Instant…

General Economics · Economics 2024-04-04 Marcus Ogren

"The chance to win given a certain move" is an easily obtainable quantity from data and often quoted in gaming statistics. It is also the fundamental quantity that reinforcement learning AI bases on. Unfortunately, this conditional…

Physics and Society · Physics 2018-03-16 I-Sheng Yang

Interleaving is an online evaluation approach for information retrieval systems that compares the effectiveness of ranking functions in interpreting the users' implicit feedback. Previous work such as Hofmann et al (2011) has evaluated the…

Information Retrieval · Computer Science 2023-03-20 Alessandro Benedetti , Anna Ruggero

Multi-round competitions often double or triple the points awarded in the final round, calling it a bonus, to maximize spectators' excitement. In a two-player competition with $n$ rounds, we aim to derive the optimal bonus size to maximize…

Computer Science and Game Theory · Computer Science 2024-06-10 Zhihuan Huang , Yuqing Kong , Tracy Xiao Liu , Grant Schoenebeck , Shengwei Xu

Energy prices and net power injection limitations regulate the operations in distribution grids and typically ensure that operational constraints are met. Nevertheless, unexpected or prolonged abnormal events could undermine the grid's…

Systems and Control · Electrical Eng. & Systems 2024-03-29 Guido Cavraro , Joshua Comden , Andrey Bernstein

In deep model compression, the recent finding "Lottery Ticket Hypothesis" (LTH) (Frankle & Carbin, 2018) pointed out that there could exist a winning ticket (i.e., a properly pruned sub-network together with original weight initialization)…

Machine Learning · Computer Science 2021-07-20 Ning Liu , Geng Yuan , Zhengping Che , Xuan Shen , Xiaolong Ma , Qing Jin , Jian Ren , Jian Tang , Sijia Liu , Yanzhi Wang

Users can now give back energies to the grid using distributed resources. Proper incentive mechanisms are required for such users, also known as prosumers, in order to maximize the sell-back amount while maintaining the retailer's profit.…

Optimization and Control · Mathematics 2022-03-14 Diptangshu Sen , Arnob Ghosh

What happens when a pretrained generative robot policy is provided a constant initial noise as input, rather than repeatedly sampling it from a Gaussian? We demonstrate that the performance of a pretrained, frozen diffusion or flow matching…

Inference-time alignment techniques offer a lightweight alternative or complement to costly reinforcement learning, while enabling continual adaptation as alignment objectives and reward targets evolve. Existing theoretical analyses justify…

Machine Learning · Computer Science 2026-05-14 Ye Wang , Jing Liu , Toshiaki Koike-Akino

Best-of-N (BoN) sampling is a widely used inference-time alignment method for language models, whereby N candidate responses are sampled from a reference model and the one with the highest predicted reward according to a learned reward…

Machine Learning · Computer Science 2026-03-09 Ved Sriraman , Adam Block

In the National Basketball Association (NBA), teams must make choices about which players to acquire, how much to pay them, and other decisions that are fundamentally dependent on player effectiveness. Thus, there is great interest in…

Applications · Statistics 2013-01-17 Dapo Omidiran

We present a hierarchical architecture to improve the efficiency of event-triggered control (ETC) in reducing resource consumption. This paper considers event-triggered systems generally as an impulsive control system in which the objective…

Systems and Control · Electrical Eng. & Systems 2024-09-17 Pio Ong , Manuel Mazo , Aaron D. Ames

Problem definition: Professional sports leagues may be suspended due to various reasons such as the recent COVID-19 pandemic. A critical question the league must address when re-opening is how to appropriately select a subset of the…

Optimization and Control · Mathematics 2024-04-02 Ali Hassanzadeh , Mojtaba Hosseini , John G. Turner

This paper explores a novel way for analyzing the tournament structures to find a best suitable one for the tournament under consideration. It concerns about three aspects such as tournament conducting cost, competitiveness development and…

Artificial Intelligence · Computer Science 2016-11-28 Nhien Pham Hoang Bao , Hiroyuki Iida

In the last round of the FIFA World Cup group stage, games for which the outcome does not affect the selection of the qualified teams are played with little enthusiasm. Furthermore, a team that has already qualified may take into account…

Applications · Statistics 2021-02-04 Mario Chater , Luc Arrondel , Jean-Pascal Gayant , Jean-François Laslier

We show that discounted methods for solving continuing reinforcement learning problems can perform significantly better if they center their rewards by subtracting out the rewards' empirical average. The improvement is substantial at…

Machine Learning · Computer Science 2024-10-31 Abhishek Naik , Yi Wan , Manan Tomar , Richard S. Sutton

Strategic decisions are often made over multiple periods of time, wherein decisions made earlier impact a competitor's success in later stages. In this paper, we study these dynamics in General Lotto games, a class of models describing the…

Computer Science and Game Theory · Computer Science 2025-05-06 Keith Paarporn , Rahul Chandan , Mahnoosh Alizadeh , Jason R. Marden

The hot-hand theory posits that an athlete who has performed well in the recent past performs better in the present. We use multilevel logistic regression to test this theory for National Hockey League playoff goaltenders, controlling for a…

Applications · Statistics 2024-05-13 Likang Ding , Ivor Cribben , Armann Ingolfsson , Monica Tran