English

LLM Inference Serving: Survey of Recent Advances and Opportunities

Distributed, Parallel, and Cluster Computing 2024-07-18 v1 Artificial Intelligence

Abstract

This survey offers a comprehensive overview of recent advancements in Large Language Model (LLM) serving systems, focusing on research since the year 2023. We specifically examine system-level enhancements that improve performance and efficiency without altering the core LLM decoding mechanisms. By selecting and reviewing high-quality papers from prestigious ML and system venues, we highlight key innovations and practical considerations for deploying and scaling LLMs in real-world production environments. This survey serves as a valuable resource for LLM practitioners seeking to stay abreast of the latest developments in this rapidly evolving field.

Keywords

Cite

@article{arxiv.2407.12391,
  title  = {LLM Inference Serving: Survey of Recent Advances and Opportunities},
  author = {Baolin Li and Yankai Jiang and Vijay Gadepally and Devesh Tiwari},
  journal= {arXiv preprint arXiv:2407.12391},
  year   = {2024}
}