中文
相关论文

相关论文: Improved Load Balancing in Large Scale Systems usi…

200 篇论文

We analyze the performance of redundancy in a multi-type job and multi-type server system. We assume the job dispatcher is unaware of the servers' capacities, and we set out to study under which circumstances redundancy improves the…

网络与互联网体系结构 · 计算机科学 2020-12-16 Elene Anton , Urtzi Ayesta , Matthieu Jonckheere , Ina Verloop

Large language model (LLM) serving is becoming an increasingly critical workload for cloud providers. Existing LLM serving systems focus on interactive requests, such as chatbots and coding assistants, with tight latency SLO requirements.…

分布式、并行与集群计算 · 计算机科学 2025-02-26 Archit Patke , Dhemath Reddy , Saurabh Jha , Haoran Qiu , Christian Pinto , Chandra Narayanaswami , Zbigniew Kalbarczyk , Ravishankar Iyer

Multiserver-job systems, where jobs require concurrent service at many servers, occur widely in practice. Essentially all of the theoretical work on multiserver-job systems focuses on maximizing utilization, with almost nothing known about…

性能 · 计算机科学 2022-11-08 Isaac Grosof , Ziv Scully , Mor Harchol-Balter , Alan Scheller-Wolf

In this paper, we consider modeling time-dependent multi-server queues that include abandonments and retrials. For the performance analysis of those, fluid and diffusion models called "strong approximations" have been widely used in the…

概率论 · 数学 2009-11-13 Young Myoung Ko , Natarajan Gautam

The model is a "generalized switch", serving multiple traffic flows in discrete time. The switch uses MaxWeight algorithm to make a service decision (scheduling choice) at each time step, which determines the probability distribution of the…

概率论 · 数学 2015-02-13 Rahul Singh , Alexander Stolyar

An increasing number of real-time applications with compute and/or communication deadlines are being supported on shared infrastructure. Such applications can often tolerate occasional deadline violations without substantially impacting…

网络与互联网体系结构 · 计算机科学 2016-03-08 Yuhuan Du , Gustavo de Veciana

To boost energy saving for the general delay-tolerant IoT networks, a two-stage and single-relay queueing communication scheme is investigated. Concretely, a traffic-aware $N$-threshold and gated-service policy are applied at the relay. As…

网络与互联网体系结构 · 计算机科学 2020-02-18 Nan Qi , Nikolaos I. Miridakis , Ming Xiao , Theodoros A. Tsiftsis , Rugui Yao , Shi Jin

The fundamental problem in the study of parallel-server systems is that of finding and analyzing `good' routing policies of arriving jobs to the servers. It is well known that, if full information regarding the workload process is available…

概率论 · 数学 2019-04-24 Pascal Moyal , Ohad Perry

Most load balancing techniques implemented in current data centers tend to rely on a mapping from packets to server IP addresses through a hash value calculated from the flow five-tuple. The hash calculation allows extremely fast packet…

网络与互联网体系结构 · 计算机科学 2017-07-11 Qingkai Liang , Sem Borst

Size-based schedulers have very desirable performance properties: optimal or near-optimal response time can be coupled with strong fairness guarantees. Despite this, such systems are very rarely implemented in practical settings, because…

分布式、并行与集群计算 · 计算机科学 2015-08-07 Matteo Dell'Amico , Damiano Carra , Pietro Michiardi

We study tandem queueing systems in which servers work more efficiently in teams than on their own and customers are impatient in that they may leave the system while waiting for service. Our goal is to determine the server assignment…

概率论 · 数学 2026-01-23 Bihan Chatterjee , Sigrún Andradóttir , Hayriye Ayhan

In this paper, we investigate scheduling policies that minimize the age of information in single-hop queueing systems. We propose a Last-Generated, First-Serve (LGFS) scheduling policy, in which the packet with the earliest generation time…

信息论 · 计算机科学 2019-04-16 Ahmed M. Bedewy , Yin Sun , Ness B. Shroff

We consider a queueing system composed of a dispatcher that routes deterministically jobs to a set of non-observable queues working in parallel. In this setting, the fundamental problem is which policy should the dispatcher implement to…

性能 · 计算机科学 2025-02-23 Jonatha Anselmi , Bruno Gaujal , Tommaso Nesti

On-Demand Ride-Pooling services have the potential to increase traffic efficiency compared to private vehicle trips by decreasing parking space needed and increasing vehicle occupancy due to higher vehicle utilization and shared trips,…

系统与控制 · 电气工程与系统科学 2023-08-11 Roman Engelhardt , Hani S. Mahmassani , Klaus Bogenberger

The load-balancing system, built on the basis of a subsystem load balancer and subsystem control and monitoring that closely interact with each other was propose in work. This system is presented as a queuing system with priority service…

网络与互联网体系结构 · 计算机科学 2019-05-10 Igor Ivanisenko , Tamara Radivilova

We develop a fluid-flow model for routing problems, where fluid consists of different size particles and the task is to route the incoming fluid to $n$ parallel servers using the size information in order to minimize the mean latency. The…

性能 · 计算机科学 2025-09-29 Runhan Xie , Esa Hyytiä , Rhonda Righter

We consider an automatic overload control for two large service systems modeled as multi-server queues, such as call centers. We assume that the two systems are designed to operate independently, but want to help each other respond to…

概率论 · 数学 2014-07-30 Ohad Perry , Ward Whitt

Shared autonomous electric vehicles can provide on-demand transportation for passengers while also interacting extensively with the electric distribution system. This interaction is especially beneficial after a disaster when the large…

系统与控制 · 电气工程与系统科学 2025-04-18 Jake Robbennolt , Meiyi Li , Javad Mohammadi , Stephen D. Boyles

Large-scale distributed computing systems often contain thousands of distributed nodes (machines). Monitoring the conditions of these nodes is important for system management purposes, which, however, can be extremely resource demanding as…

分布式、并行与集群计算 · 计算机科学 2019-05-23 Tiffany Tuor , Shiqiang Wang , Kin K. Leung , Bong Jun Ko

We study a job-assignment problem in a large-scale server farm system with geographically deployed servers as abstracted computer components (e.g., storage, network links, and processors) that are potentially diverse. We aim to maximize the…

分布式、并行与集群计算 · 计算机科学 2020-03-30 Jing Fu , Bill Moran