Related papers: Logarithmic Heavy Traffic Error Bounds in Generali…
The deployment of machine learning in high-stakes services relies on ``human-in-the-loop'' architectures to mitigate algorithmic uncertainty. However, existing static policies fail to address a fundamental tension: algorithms suffer from…
Over 40% of computational power in Large Language Model (LLM) serving systems can be systematically wasted - not from hardware limits, but from load imbalance in barrier-synchronized parallel processing. When progress is gated by the…
We investigate the functional limits of generalized Jackson networks in a multi-scale heavy traffic regime where stations approach full utilization at distinct, separated rates. Our main result shows that the appropriately scaled queue…
A fundamental challenge in large-scale cloud networks and data centers is to achieve highly efficient server utilization and limit energy consumption, while providing excellent user-perceived performance in the presence of uncertain and…
We examine a queue-based random-access algorithm where activation and deactivation rates are adapted as functions of queue lengths. We establish its heavy traffic behavior on a complete interference graph, which turns out to be highly…
We study a single server FIFO queue that offers general service. Each of n customers enter the queue at random time epochs that are inde- pendent and identically distributed. We call this the random scattering traffic model, and the…
In this paper, we study stability and latency of routing in wireless networks where it is assumed that no collision will occur. Our approach is inspired by the adversarial queuing theory, which is amended in order to model wireless…
Heavy traffic analysis for load balancing policies has relied heavily on the condition of state-space collapse onto a single-dimensional line in previous works. In this paper, via Lyapunov-drift analysis, we rigorously prove that even under…
In the online packet buffering problem (also known as the unweighted FIFO variant of buffer management), we focus on a single network packet switching device with several input ports and one output port. This device forwards unit-size,…
Packet traffic in complex networks undergoes the jamming transition from free-flow to congested state as the number of packets in the system increases. Here we study such jamming transition when queues are operated by the priority queuing…
We consider the heavy-traffic approximation to the $GI/M/s$ queueing system in the Halfin-Whitt regime, where both the number of servers $s$ and the arrival rate $\lambda$ grow large (taking the service rate as unity), with…
We analyze opportunistic schemes for transmission scheduling from one of $n$ homogeneous queues whose channel states fluctuate independently. Considered schemes consist of the LCQ policy, which transmits from a longest connected queue in…
We present an overview of scalable load balancing algorithms which provide favorable delay performance in large-scale systems, and yet only require minimal implementation overhead. Aimed at a broad audience, the paper starts with an…
With the rapid advance of information technology, network systems have become increasingly complex and hence the underlying system dynamics are often unknown or difficult to characterize. Finding a good network control policy is of…
Modern networks exhibit a high degree of variability in link rates. Cellular network bandwidth inherently varies with receiver motion and orientation, while class-based packet scheduling in datacenter and service provider networks induces…
In this thesis, we propose and analyze a multi-server model that captures a performance trade-off between centralized and distributed processing. In our model, a fraction $p$ of an available resource is deployed in a centralized manner…
In this work, we study the stationary distribution of the scaled queue length vector process in multiclass queueing networks operating under static buffer priority service policies. We establish that when subjected to a multi-scale heavy…
The Join-the-Shortest-Queue (JSQ) load-balancing scheme is known to minimise the average delay of jobs in homogeneous systems consisting of identical servers. However, it performs poorly in heterogeneous systems where servers have different…
We establish a unified analytical framework for load balancing systems, which allows us to construct a general class $\Pi$ of policies that are both throughput optimal and heavy-traffic delay optimal. This general class $\Pi$ includes as…
We consider the FCFS $GI/GI/n$ queue in the Halfin-Whitt heavy traffic regime, and prove bounds for the steady-state probability of delay (s.s.p.d.) for generally distributed processing times. We prove that there exist $\epsilon_1,…