Related papers: Fluid Limits for Time-Varying Many-Server Queues w…
In this paper, we analyze a discrete-time queue that is motivated from studying hospital inpatient flow management, where the customer count process captures the midnight inpatient census. The stationary distribution of the customer count…
We study infinite server queues driven by Cox processes in a fast oscillatory random environment. While exact performance analysis is difficult, we establish diffusion approximations to the (re-scaled) number-in-system process by proving…
The Foster-Lyapunov theorem and its variants serve as the primary tools for studying the stability of queueing systems. In addition, it is well known that setting the drift of the Lyapunov function equal to zero in steady-state provides…
Although density functional theory provides reliable predictions for the static properties of simple fluids under confinement, a theory of comparative accuracy for the transport coefficients has yet to emerge. Nonetheless, there is evidence…
The Quality-and-Efficiency-Driven (QED) regime provides a basis for solving asymptotic dimensioning problems that trade off revenue, costs and service quality. We derive bounds for the optimality gaps that capture the differences between…
Arrival processes to service systems often display (i) larger than anticipated fluctuations, (ii) a time-varying rate, and (iii) temporal correlation. Motivated by this, we introduce a specific non-homogeneous Poisson process that…
In this paper, we study a large system of $N$ servers each with capacity to process at most $C$ simultaneous jobs and an incoming job is routed to a server if it has the lowest occupancy amongst $d$ (out of N) randomly selected servers. A…
The Join-the-Shortest-Queue (JSQ) load-balancing scheme is known to minimise the average delay of jobs in homogeneous systems consisting of identical servers. However, it performs poorly in heterogeneous systems where servers have different…
In this paper, we consider a L\'evy-driven fluid queueing system where the server may subject to breakdowns and repairs. In addition, the server will leave for a vacation each time when he finds an empty system. We cast the queueing process…
We study a multiclass M/M/1 queueing control problem with finite buffers under heavy-traffic where the decision maker is uncertain about the rates of arrivals and service of the system and by scheduling and admission/rejection decisions…
We present an analysis of large-scale load balancing systems, where the processing time distribution of tasks depends on both the task and server types. Our study focuses on the asymptotic regime, where the number of servers and task types…
A multi-class single-server queueing model with finite buffers, in which scheduling and admission of customers are subject to control, is studied in the moderate deviation heavy traffic regime. A risk-sensitive cost set over a finite time…
We consider the numerical approximation of compressible flow in a pipe network. Appropriate coupling conditions are formulated that allow us to derive a variational characterization of solutions and to prove global balance laws for the…
We study a simple rate control scheme for a multiclass queuing network for which customers are partitioned into distinct flows that are queued separately at each station. The control scheme discards customers that arrive to the network…
We construct the hydrodynamic theory of coherent collective motion ("flocking") at a solid-liquid interface. The polar order parameter and concentration of a collection of "active" (self-propelled) particles at a planar interface between a…
In this paper, we focus our attention on the large capacities unsplittable flow problem in a game theoretic setting. In this setting, there are selfish agents, which control some of the requests characteristics, and may be dishonest about…
This paper addresses the analysis of the queue-length process of single-server queues under overdispersion, i.e., queues fed by an arrival process for which the variance of the number of arrivals in a given time window exceeds the…
Large language models now serve millions of users daily, with providers incurring costs exceeding $700,000 per day. Each request requires token-by-token inference, making GPU scheduling central to latency, capacity, and cost. The difficulty…
We consider large-scale service systems with multiple customer classes and multiple server pools; interarrival and service times are exponentially distributed, and mean service times depend both on the customer class and server pool. It is…
Discrete-time queueing system has widespread applications in packet switching networks, internet protocol, Broadband Integrated Services Digital Network (B-ISDN), circuit switched time-division multiple access etc. In this paper, we analyze…