Related papers: Achieving Maximum Utilization in Optimal Time for …
A novel phase transition behaviour is observed in the Kolkata Paise Restaurant (KPR) problem where large number ($N$) of agents or customers collectively (and iteratively) learn to choose among the $N$ restaurants where she would expect to…
We study the dynamics of a few stochastic learning strategies for the 'Kolkata Paise Restaurant' problem, where N agents choose among N equally priced but differently ranked restaurants every evening such that each agent tries get to dinner…
We study the dynamics of some uniform learning strategy limits or a probabilistic version of the "Kolkata Paise Restaurant" problem, where N agents choose among N equally priced but differently ranked restaurants every evening such that…
We study the dynamics of the "Kolkata Paise Restaurant problem". The problem is the following: In each period, N agents have to choose between N restaurants. Agents have a common ranking of the restaurants. Restaurants can only serve one…
We will review the results for stochastic learning strategies, both classical (one-shot and iterative) and quantum (one-shot only), for optimizing the available many-choice resources among a large number of competing agents, developed over…
We study the Kolkata Paise Restaurant Problem (KPRP) with multiple dining clubs, extending work in [A. Harlalka, A. Belmonte and C. Griffin, \textit{Physica A}, 620:128767, 2023]. In classical KPRP, $N$ agents chose among $N$ restaurants at…
In this paper, we study a large-scale distributed coordination problem and propose efficient adaptive strategies to solve the problem. The basic problem is to allocate finite number of resources to individual agents such that there is as…
In the Constrained Fault-Tolerant Resource Allocation (FTRA) problem, we are given a set of sites containing facilities as resources, and a set of clients accessing these resources. Specifically, each site i is allowed to open at most R_i…
We discuss the strategy that rational agents can use to maximize their expected long-term payoff in the co-action minority game. We argue that the agents will try to get into a cyclic state, where each of the $(2N +1)$ agent wins exactly…
This paper investigates the application of Reinforcement Learning (RL) to optimise call routing in call centres to minimise client waiting time and staff idle time. Two methods are compared: a model-based approach using Value Iteration (VI)…
Large language models trained with reinforcement learning with verifiable rewards tend to trade accuracy for length--inflating response lengths to achieve gains in accuracy. While longer answers may be warranted for harder problems, many…
First-price auctions have largely replaced traditional bidding approaches based on Vickrey auctions in programmatic advertising. As far as learning is concerned, first-price auctions are more challenging because the optimal bidding strategy…
Model-free Reinforcement Learning (RL) generally suffers from poor sample complexity, mostly due to the need to exhaustively explore the state-action space to find well-performing policies. On the other hand, we postulate that expert…
When deploying artificial agents in real-world environments where they interact with humans, it is crucial that their behavior is aligned with the values, social norms or other requirements of that environment. However, many environments…
The present study proposes clustering techniques for designing demand response (DR) programs for commercial and residential prosumers. The goal is to alter the consumption behavior of the prosumers within a distributed energy community in…
The route planning problem based on the greedy algorithm represents a method of identifying the optimal or near-optimal route between a given start point and end point. In this paper, the PCA method is employed initially to downscale the…
In this article, we present a brief narration of the origin and the overview of the recent developments done on the Kolkata Paise Restaurant (KPR) problem, which can serve as a prototype for a broader class of resource allocation problems…
We consider an infinite collection of agents who make decisions, sequentially, about an unknown underlying binary state of the world. Each agent, prior to making a decision, receives an independent private signal whose distribution depends…
Autonomous agents must often deal with conflicting requirements, such as completing tasks using the least amount of time/energy, learning multiple tasks, or dealing with multiple opponents. In the context of reinforcement learning~(RL),…
We consider the Monte-Carlo first visit algorithm, of which the goal is to find the optimal control in a Markov decision process with finite state space and finite number of possible actions. We show its convergence when the discount factor…