Related papers: On Sojourn Times in the Finite Capacity $M/M/1$ Qu…
This paper focuses on an infinite-server queue modulated by an independently evolving finite-state Markovian background process, with transition rate matrix $Q\equiv(q_{ij})_{i,j=1}^d$. Both arrival rates and service rates are depending on…
We study a sequence of single server queues with customer abandonment (GI/GI/1+GI) under heavy traffic. The patience time distributions vary with the sequence, which allows for a wider scope of applications. It is known ([20, 18]) that the…
Recently, the problem of multitasking scheduling has attracted a lot of attention in the service industries where workers frequently perform multiple tasks by switching from one task to another. Hall, Leung and Li (Discrete Applied…
Imagine, you enter a grocery store to buy food. How many peopledo you overlap with in this store? How much time do you overlap witheach person in the store? In this paper, we answer these questions bystudying the overlap times between…
The Join-the-Shortest-Queue-d routing policy is considered for a large system with $n$ servers. Moderate deviation principles (MDP) for the occupancy process and the empirical queue length process are established as $n\to \infty$. Each MDP…
Ride-pooling systems, despite being an appealing urban mobility mode, still struggle to gain momentum. While we know the significance of critical mass in reaching system sustainability, less is known about the spatiotemporal patterns of…
Distributions with a heavy tail are difficult to estimate. If the design of a scheduling policy is sensitive to the details of heavy tail distributions of the service times, an approximately optimal solution is difficult to obtain. This…
This is an expository review paper illustrating the ``martingale method'' for proving many-server heavy-traffic stochastic-process limits for queueing models, supporting diffusion-process approximations. Careful treatment is given to an…
In discrete time, customers arrive at random. Each waits until one of three servers is available; each thereafter departs at random. We seek the distribution of maximum line length of idle customers. Algebraic expressions obtained for the…
We consider a setting where qubits are processed sequentially, and derive fundamental limits on the rate at which classical information can be transmitted using quantum states that decohere in time. Specifically, we model the sequential…
We study a spatiotemporal service matching problem in which demand, heterogeneous in location and time sensitivity/preference, is to be assigned to service stations. The planner seeks to maximize social welfare, defined as total service…
We consider a single large language model (LLM) server that serves a heterogeneous stream of queries belonging to $N$ distinct task types. Queries arrive according to a Poisson process, and each type occurs with a known prior probability.…
This paper considers the dispatching of large-scale real-time ride-sharing systems to address congestion issues faced by many cities. The goal is to serve all customers (service guarantees) with a small number of vehicles while minimizing…
We obtain asymptotic bounds for the tail distribution of steady-state waiting time in a two server queue where each server processes incoming jobs at a rate equal to the rate of their arrivals (that is, the half-loaded regime). The job…
In this paper we consider the problem of maximum throughput for tandem queueing system. We modeled this system as a Quasi-Birth-Death process. In order to do this we named level the number of customers waiting in the first buffer (including…
We consider a two-queue polling model with switch-over times and $k$-limited service (serve at most $k_i$ customers during one visit period to queue $i$) in each queue. The major benefit of the $k$-limited service discipline is that it -…
Problem Definition: Allocating sufficient capacity to cloud services is a challenging task, especially when demand is time-varying, heterogeneous, contains batches, and requires multiple types of resources for processing. In this setting,…
We present a formalism to study many-particle quantum transport across a lattice locally connected to two finite, non-stationary (bosonic or fermionic) reservoirs, both of which are in a thermal state. We show that, for conserved total…
We consider a discrete-time parallel service system consisting of $n$ heterogeneous single server queues with infinite capacity. Jobs arrive to the system as an i.i.d. process with rate proportional to $n$, and must be immediately…
We study a single server queue under a processor-sharing type of scheduling policy, where the weights for determining the sharing are given by functions of each job's remaining service(processing) amount, and obtain a fluid limit for the…