English
Related papers

Related papers: SkyServe: Serving AI Models across Regions and Clo…

200 papers

The high computational and memory requirements of generative large language models (LLMs) make it challenging to serve them cheaply. This paper aims to reduce the monetary cost for serving LLMs by leveraging preemptible GPU instances on…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-11-28 Xupeng Miao , Chunan Shi , Jiangfei Duan , Xiaoli Xi , Dahua Lin , Bin Cui , Zhihao Jia

AI batch jobs such as model training, inference pipelines, and data analytics require substantial GPU resources and often need to finish before a deadline. Spot instances offer 3-10x lower cost than on-demand instances, but their…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-13 Zhifei Li , Tian Xia , Ziming Mao , Zihan Zhou , Ethan J. Jackson , Jamison Kerney , Zhanghao Wu , Pratik Mishra , Yi Xu , Yifan Qiao , Scott Shenker , Ion Stoica

Microservices architecture, known for its agility and efficiency, is an ideal framework for cloud-based software development and deployment. When integrated with containerization and orchestration systems, resource management becomes more…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-02-10 Dasith Edirisinghe , Kavinda Rajapakse , Pasindu Abeysinghe , Sunimal Rathnayake

Cloud service platforms increasingly rely on elastic infrastructures to support dynamic workloads. Spot instances provide discounted computing resources but introduce uncertainty due to dynamic pricing, resource availability, and…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-22 Javier Fabra , Enrique Molina-Giménez , Pedro García-López

Cloud vendors offer discounted spot instances to maximize surplus resource utilization, but these instances are subject to the risk of sudden interruption. Traditional pricing datasets have been employed to predict this risk, yet recent…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-28 Taeyoon Kim , Kyumin Kim , Kyunghwan Kim , Hayoung Kim , Seungwoo Jeong , Moohyun Song , Kyungyong Lee

Public cloud service vendors provide a surplus of computing resources at a cheaper price as a spot instance. Despite the cheaper price, the spot instance can be forced to be shutdown at any moment whenever the surplus resources are in…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-10-26 Sungjae Lee , Jaeil Hwang , Kyungyong Lee

Cloud providers sell their idle capacity on markets through an auction-like mechanism to increase their return on investment. The instances sold in this way are called spot instances. In spite that spot instances are usually 90% cheaper…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-03-08 Chenhao Qu , Rodrigo N. Calheiros , Rajkumar Buyya

Cloud computing providers are now offering their unused resources for leasing in the spot market, which has been considered the first step towards a full-fledged market economy for computational resources. Spot instances are virtual…

Distributed, Parallel, and Cluster Computing · Computer Science 2011-10-28 William Voorsluys , Rajkumar Buyya

The demand for smartness in embedded systems has been mounting up drastically in the past few years. Embedded system today must address the fundamental challenges introduced by cloud computing and artificial intelligence. KubeEdge [1] is an…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-07-21 Sean Wang , Yuxiao Hu , Jason Wu

The surging development of Artificial Intelligence-Generated Content (AIGC) marks a transformative era of the content creation and production. Edge servers promise attractive benefits, e.g., reduced service delay and backhaul traffic load,…

Machine Learning · Computer Science 2024-09-10 Yuxin Liang , Peng Yang , Yuanyuan He , Feng Lyu

Cost optimization is a common goal of workflow schedulers operating in cloud computing environments. The use of spot instances is a potential means of achieving this goal, as they are offered by cloud providers at discounted prices compared…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-08-07 Amanda Jayanetti , Saman Halgamuge , Rajkumar Buyya

Modern applications span multiple clouds to reduce costs, avoid vendor lock-in, and leverage low-availability resources in another cloud. However, standard object stores operate within a single cloud, forcing users to manually manage data…

Recent developments in large language models (LLMs) have demonstrated their remarkable proficiency in a range of tasks. Compared to in-house homogeneous GPU clusters, deploying LLMs in cloud environments with diverse types of GPUs is…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-11-07 Youhe Jiang , Fangcheng Fu , Xiaozhe Yao , Taiyi Wang , Bin Cui , Ana Klimovic , Eiko Yoneki

Large language models (LLMs) power many modern applications, but serving them at scale remains costly and resource-intensive. Current server-centric systems overlook consumer-grade GPUs at the edge. We introduce SpecEdge, an edge-assisted…

Computation and Language · Computer Science 2025-11-19 Jinwoo Park , Seunggeun Cho , Dongsu Han

Modern applications increasingly rely on inference serving systems to provide low-latency insights with a diverse set of machine learning models. Existing systems often utilize resource elasticity to scale with demand. However, many…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-05-13 Joel Wolfrath , Daniel Frink , Abhishek Chandra

Cloud users aim to minimize cost while maximizing performance by selecting the most suitable instance types for their workloads. To reduce expenses, spot instances have been widely adopted due to their steep discounts compared to on-demand…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-28 Taeyoon Kim , Kyumin Kim , Enrique Molina-Giménez , Pedro García-López , Kyungyong Lee

Generative recommendation (GR) offers superior modeling capabilities but suffers from prohibitive inference costs due to the repeated encoding of long user histories. While cross-request Key-Value (KV) cache reuse presents a significant…

We propose SparsePipe, an efficient and asynchronous parallelism approach for handling 3D point clouds with multi-GPU training. SparsePipe is built to support 3D sparse data such as point clouds. It achieves this by adopting generalized…

Computer Vision and Pattern Recognition · Computer Science 2020-12-29 Keke Zhai , Pan He , Tania Banerjee , Anand Rangarajan , Sanjay Ranka

To meet next-generation IoT application demands, edge computing moves processing power and storage closer to the network edge to minimise latency and bandwidth utilisation. Edge computing is becoming popular as a result of these benefits,…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-12-13 Aadharsh Roshan Nandhakumar , Ayush Baranwal , Priyanshukumar Choudhary , Muhammed Golec , Sukhpal Singh Gill

Shared edge computing platforms deployed at the radio access network are expected to significantly improve quality of service delivered by Application Service Providers (ASPs) in a flexible and economic way. However, placing edge service in…

Networking and Internet Architecture · Computer Science 2018-10-09 Lixing Chen , Jie Xu , Shaolei Ren , Pan Zhou
‹ Prev 1 2 3 10 Next ›