Related papers: SPLIT: SymPathy for Large jobs Improves Tail laten…
In the literature, retrial queues with batch arrivals and heavy service times have been studied and the so-called equivalence theorem has been established under the condition that the service time is heavier than the batch size. The…
We consider the problem of scheduling a queueing system in which many statistically identical servers cater to several classes of impatient customers. Service times and impatience clocks are exponential while arrival processes are renewal.…
The Shortest Remaining Processing Time (SRPT) scheduling policy and its variants have been extensively studied in both theoretical and practical settings. While beautiful results are known for single-server SRPT, much less is known for…
Reinforcement Learning (RL) is a pivotal post-training technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, synchronous RL post-training often suffers from significant GPU underutilization, referred…
Two of the most popular approximations for the distribution of the steady-state waiting time, $W_{\infty}$, of the M/G/1 queue are the so-called heavy-traffic approximation and heavy-tailed asymptotic, respectively. If the traffic…
We consider a two-node fluid network with batch arrivals of random size having a heavy-tailed distribution. We are interested in the tail asymptotics for the stationary distribution of a two-dimensional queue-length process. The tail…
Modern data centers serve workloads which are capable of exploiting parallelism. When a job parallelizes across multiple servers it will complete more quickly, but jobs receive diminishing returns from being allocated additional servers.…
In this paper, by the singular-perturbation technique, we investigate the heavy-traffic behavior of a priority polling system consisting of three M/M/1 queues with threshold policy. It turns out that the scaled queue-length of the…
Modern latency-critical online services often rely on composing results from a large number of server components. Hence the tail latency (e.g. the 99th percentile of response time), rather than the average, of these components determines…
Significant correlations between arrivals of load-generating events make the numerical evaluation of the workload of a system a challenging problem. In this paper, we construct highly accurate approximations of the workload distribution of…
How should we schedule jobs to minimize mean queue length? In the preemptive M/G/1 queue, we know the optimal policy is the Gittins policy, which uses any available information about jobs' remaining service times to dynamically prioritize…
Multiserver jobs, which are jobs that occupy multiple servers simultaneously during service, are prevalent in today's computing clusters. But little is known about the delay performance of systems with multiserver jobs. We consider queueing…
In this paper, we study the maximum waiting time $\max_{i\leq N}W_i(\cdot)$ in an $N$-server fork-join queue with heavy-tailed services as $N\to\infty$. The service times are the product of two random variables. One random variable has a…
In this paper, we consider a discrete-time preemptive priority queue with different service rates for two classes of customers, one with high-priority and the other with low-priority. This model corresponds to the classical preemptive…
The imbalance (or long-tail) is the nature of many real-world data distributions, which often induces the undesirable bias of deep classification models toward frequent classes, resulting in poor performance for tail classes. In this paper,…
We study the asymptotics of the stationary sojourn time Z of a "typical customer" in a tandem of single-server queues. It is shown that, in a certain "intermediate" region of light-tailed service time distributions, Z may take a large value…
The areas under workload process and under queuing process in a single server queue over the busy period have many applications not only in queuing theory but also in risk theory or percolation theory. We focus here on the tail behaviour of…
In the context of communication networks, the framework of stochastic event graphs allows a modeling of control mechanisms induced by the communication protocol and an analysis of its performances. We concentrate on the logarithmic tail…
When an explicit expression for a probability distribution function $F(x)$ can not be found, asymptotic properties of the tail probability function $\bar{F}(x)=1-F(x)$ are very valuable, since they provide approximations or bounds for…
Large-scale conversational systems typically rely on a skill-routing component to route a user request to an appropriate skill and interpretation to serve the request. In such system, the agent is responsible for serving thousands of skills…