相关论文: The M/M/Infinity Service System with Ranked Server…
In Large Language Model (LLM) inference, the output length of an LLM request is typically regarded as not known a priori. Consequently, most LLM serving systems employ a simple First-come-first-serve (FCFS) scheduling strategy, leading to…
The paper studies a closed queueing network containing two types of node. The first type (server station) is an infinite server queueing system, and the second type (client station) is a single server queueing system with autonomous…
We consider a distributed server system consisting of a large number of servers, each with limited capacity on multiple resources (CPU, memory, disk, etc.). Jobs with different rewards arrive over time and require certain amounts of…
The paper studies closed queueing networks containing a server station and $k$ client stations. The server station is an infinite server queueing system, and client stations are single-server queueing systems with autonomous service, i.e.…
In this paper we consider an M/G/1-type queue fed by a finite customer-pool. In terms of transforms, we characterize the time-dependent distribution of the number of customers and the workload, as well as the associated waiting times.
Queues that feature multiple entities arriving simultaneously are among the oldest models in queueing theory, and are often referred to as "batch" (or, in some cases, "bulk") arrival queueing systems. In this work we study the affect of…
In this paper, we consider a two server system serving heterogeneous customers. One of the server has a FIFO scheduling policy and charges a fixed admission price to each customer. The second queue follows the highest-bidder-first (HBF)…
We investigate the scheduling of a common resource between several concurrent users when the feasible transmission rate of each user varies randomly over time. Time is slotted and users arrive and depart upon service completion. This may…
This paper addresses the analysis of the queue-length process of single-server queues under overdispersion, i.e., queues fed by an arrival process for which the variance of the number of arrivals in a given time window exceeds the…
We study a fundamental model of resource allocation in which a finite number of resources must be assigned in an online manner to a heterogeneous stream of customers. The customers arrive randomly over time according to known stochastic…
We consider a single-server system with service stations in each point of the circle. Customers arrive after exponential times at uniformly-distributed locations. The server moves at finite speed and adopts a greedy routing mechanism. It…
We study a double-ended queue which consists of two classes of customers. Whenever there is a pair of customers from both classes, they are matched and leave the system immediately. The matching follows first-come-first-serve principle. If…
This paper concerns the recurrence structure of the infinite server queue, as viewed through the prism of the maximum dater sequence, namely the time to drain the current work in the system as seen at arrival epochs. Despite the importance…
There are $n$ queues, each with a single server. Customers arrive in a Poisson process at rate $\lambda n$, where $0<\lambda<1$. Upon arrival each customer selects $d\geq2$ servers uniformly at random, and joins the queue at a least-loaded…
This paper studies the heavy-traffic (HT) behaviour of queueing networks with a single roving server. External customers arrive at the queues according to independent renewal processes and after completing service, a customer either leaves…
A single server retrial queueing system with two-classes of orbiting customers, and general class dependent service times is considered. If an arriving customer finds the server unavailable, it enters a virtual queue, called the orbit,…
Network traffic analysis increasingly uses complex machine learning models as the internet consolidates and traffic gets more encrypted. However, over high-bandwidth networks, flows can easily arrive faster than model inference rates. The…
This paper considers a parallel system of queues fed by independent arrival streams, where the service rate of each queue depends on the number of customers in all of the queues. Necessary and sufficient conditions for the stability of the…
Recurrence and ergodic properties are established for a single--server queueing system with variable intensities of arrivals and service. Convergence to stationarity is also interpreted in terms of reliability theory.
We study a many-server queuing system with general service time distribution and state dependent service rates. The dynamics of the system are modeled using measure valued processes which keep track of the residual service times. Under…