相关论文: A Many-Server Functional Strong Law For A Non-Stat…
With the development of edge networks and mobile computing, the need to serve heterogeneous data sources at the network edge requires the design of new distributed machine learning mechanisms. As a prevalent approach, Federated Learning…
In previous papers we developed a deterministic fluid approximation for an overloaded Markovian queueing system having two customer classes and two service pools, known in the call-center literature as the X model. The system uses the…
Large language models (LLMs) have surged in popularity and are extensively used in commercial applications, where the efficiency of model serving is crucial for the user experience. Most current research focuses on optimizing individual…
In this thesis, we propose and analyze a multi-server model that captures a performance trade-off between centralized and distributed processing. In our model, a fraction $p$ of an available resource is deployed in a centralized manner…
Our goal in this paper is to investigate the fluid picture associated with an open large scale storage network of non-reliable file servers with finite capacity. In this storage system, new files can be added and a file with only one copy…
Large Language Model (LLM) serving systems remain fundamentally fragile, where frequent hardware faults in hyperscale clusters trigger disproportionate service outages in the software stack. Current recovery mechanisms are prohibitively…
In this paper, we present a functional fluid limit theorem and a functional central limit theorem for a queue with an infinity of servers M/GI/$\infty$. The system is represented by a point-measure valued process keeping track of the…
We consider N single server infinite buffer queues with service rate \beta. Customers arrive at rate N\alpha, choose L queues uniformly, and join the shortest. We study the processes R^N for large N, where R^N_t(k) is the fraction of queues…
Consider a system of identical server pools where tasks with exponentially distributed service times arrive as a time-inhomogenenous Poisson process. An admission threshold is used in an inner control loop to assign incoming tasks to server…
We obtain the law of large numbers (LLN) and the central limit theorem (CLT) for weakly dependent non-stationary arrays of random fields with asymptotically unbounded moments. The weak dependence condition for arrays of random fields is…
We study a spherical, self-gravitating fluid model, which finds applications in cosmic structure formation. We argue that since the system features nonlinearity and gravity-induced dispersion, the emergence of solitons becomes possible. We…
A nonlinear model relating the imposed motion of a circular cylinder, submerged in a fluid flow, to the transverse force coefficient is presented. The nonlinear fluid system, featuring vortex shedding patterns, limit cycle oscillations and…
In this paper we establish a weak and a strong law of large numbers for supercritical superprocesses with general non-local branching mechanisms. Our results complement earlier results obtained for superprocesses with only local branching.…
The aim of this paper is to prove the strong law of large numbers (SLLN) as well as the central limit theorem (CLT) for a class of vector-valued stochastic processes which arise as solutions of the stochastic evolution inclusion…
We investigate, under general stationary ergodic assumptions, the stability of systems of $S$ parallel queues in which any incoming customer joins the queue of the server having the $p+1$-th shortest workload ($p < S$), or a free server if…
This work considers a many-server queueing system in which impatient customers with i.i.d., generally distributed service times and i.i.d., generally distributed patience times enter service in the order of arrival and abandon the queue if…
This paper studies a single server queue in heavy traffic, with general inter-arrival and service time distributions, where arrival and service rates vary discontinuously as a function of the (diffusively scaled) queue length. It is proved…
Large language model (LLM) serving demands low latency and high throughput, but high load variability makes it challenging to achieve high GPU utilization. In this paper, we identify a synergetic but overlooked opportunity to co-serve…
Serverless computing has emerged as a compelling solution for cloud-based model inference. However, as modern large language models (LLMs) continue to grow in size, existing serverless platforms often face substantial model startup…
When predicting scalar responses in the situation where the explanatory variables are functions, it is sometimes the case that some functional variables are related to responses linearly while other variables have more complicated…