English
Related papers

Related papers: How Far Can Client-Only Solutions Go for Mobile Br…

200 papers

Session length is a very important aspect in determining a user's satisfaction with a media streaming service. Being able to predict how long a session will last can be of great use for various downstream tasks, such as recommendations and…

Information Retrieval · Computer Science 2017-08-02 Theodore Vasiloudis , Hossein Vahabi , Ross Kravitz , Valery Rashkov

Recent advancements in speculative decoding have demonstrated considerable speedup across a wide array of large language model (LLM) tasks. Speculative decoding inherently relies on sacrificing extra memory allocations to generate several…

Machine Learning · Computer Science 2025-06-04 Selin Yildirim , Deming Chen

Speed scaling for a tandem server setting is considered, where there is a series of servers, and each job has to be processed by each of the servers in sequence. Servers have a variable speed, their power consumption being a convex…

Data Structures and Algorithms · Computer Science 2019-07-11 Rahul Vaze , Jayakrishnan Nair

Serverless computing relieves developers from the burden of resource management, thus providing ease-of-use to the users and the opportunity to optimize resource utilization for the providers. However, today's serverless systems lack…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-01-26 Prasoon Sinha , Kostis Kaffes , Neeraja J. Yadwadkar

Caching popular contents at edge devices is an effective solution to alleviate the burden of the backhaul networks. Earlier investigations commonly neglected the storage cost in caching. More recently, retention-aware caching, where both…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-01-10 Ghafour Ahani , Di Yuan

Recent studies show that the coded caching technique can facilitate the wireless content distribution by mitigating the wireless traffic rate during the peak-traffic time, where the contents are partially prefetched to the local cache of…

Networking and Internet Architecture · Computer Science 2014-07-08 Sinong Wang , Xiaohua Tian , Hui Liu

A new form of caching, namely application-level caching, has been recently employed in web applications to improve their performance and increase scalability. It consists of the insertion of caching logic into the application base code to…

Software Engineering · Computer Science 2020-11-03 Jhonny Mertz , Ingrid Nunes

While mobile edge computing (MEC) alleviates the computation and power limitations of mobile devices, additional latency is incurred when offloading tasks to remote MEC servers. In this work, the power-delay tradeoff in the context of task…

Networking and Internet Architecture · Computer Science 2017-10-03 Chen-Feng Liu , Mehdi Bennis , H. Vincent Poor

Thanks to the abundance of Web platforms and broadband connections, HTTP Adaptive Streaming has become the de facto choice for multimedia delivery nowadays. However, the visual quality of adaptive video streaming may fluctuate strongly…

Multimedia · Computer Science 2020-02-26 Huyen T. T. Tran , Nam Pham Ngoc , Tobias Hoßfeld , Michael Seufert , Truong Cong Thang

Given the rapid rise in energy demand by data centers and computing systems in general, it is fundamental to incorporate energy considerations when designing (scheduling) algorithms. Machine learning can be a useful approach in practice by…

Data Structures and Algorithms · Computer Science 2021-12-07 Antonios Antoniadis , Peyman Jabbarzade Ganje , Golnoosh Shahkarami

By year 2020, the number of smartphone users globally will reach 3 Billion and the mobile data traffic (cellular + WiFi) will exceed PC internet traffic the first time. As the number of smartphone users and the amount of data transferred…

Networking and Internet Architecture · Computer Science 2017-07-24 Kemal Guner , Tevfik Kosar

Redundant transfer of resources is a critical issue for compromising the performance of mobile Web applications (a.k.a., apps) in terms of data traffic, load time, and even energy consumption. Evidence shows that the current cache…

Software Engineering · Computer Science 2016-12-12 Xuanzhe Liu , Yun Ma , Shuailiang Dong , Yunxin Liu , Tao Xie , Gang Huang , Hong Mei

Mobile data traffic (cellular + WiFi) will exceed PC Internet traffic by 2020. As the number of smartphone users and the amount of data transferred per smartphone grow exponentially, limited battery power is becoming an increasingly…

Networking and Internet Architecture · Computer Science 2018-05-23 Kemal Guner , MD S Q Zulkar Nine , Tevfik Kosar , Fatih Bulut

The growth in the number of parameters of Large Language Models (LLMs) has led to a significant surge in computational requirements, making them challenging and costly to deploy. Speculative decoding (SD) leverages smaller models to…

Computation and Language · Computer Science 2025-04-04 Matthieu Zimmer , Milan Gritta , Gerasimos Lampouras , Haitham Bou Ammar , Jun Wang

Harnessing information about the user mobility pattern and daily demand can enhance the network capability to improve the quality of experience (QoE) at Vehicular Ad-Hoc Networks (VANETs). Proactive caching, as one of the key features…

Networking and Internet Architecture · Computer Science 2018-10-03 Yousef AlNagar , Sameh Hosny , Amr A. El-Sherif

Our paper presents solutions that can significantly improve the delay performance of putting and retrieving data in and out of cloud storage. We first focus on measuring the delay performance of a very popular cloud storage service Amazon…

Networking and Internet Architecture · Computer Science 2013-11-04 Guanfeng Liang , Ulas C. Kozat

This paper analyzes the use of variable speed limits to optimize travel time reliability for commuters. The investigation focuses on a traffic corridor with a bottleneck subject to the capacity drop phenomenon. The optimization criterion is…

Optimization and Control · Mathematics 2025-09-16 Alexander Hammerl , Ravi Seshadri , Thomas Kjær Rasmussen , Otto Anker Nielsen

Speculative decoding is a pivotal technique to accelerate the inference of large language models (LLMs) by employing a smaller draft model to predict the target model's outputs. However, its efficacy can be limited due to the low predictive…

Artificial Intelligence · Computer Science 2024-06-11 Xiaoxuan Liu , Lanxiang Hu , Peter Bailis , Alvin Cheung , Zhijie Deng , Ion Stoica , Hao Zhang

Mobile applications have become an inseparable part of people's daily life. Nonetheless, the market competition is extremely fierce, and apps lacking recognition among most users are susceptible to market elimination. To this end,…

Software Engineering · Computer Science 2023-11-28 Shiqi Duan , Jianxun Liu , Yong Xiao , Xiangping Zhang

Inference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from CPU speculative execution. These techniques reduce…

Computation and Language · Computer Science 2024-11-19 Branden Butler , Sixing Yu , Arya Mazaheri , Ali Jannesari