English

A Review on Edge Large Language Models: Design, Execution, and Applications

Distributed, Parallel, and Cluster Computing 2025-02-25 v2

Abstract

Large language models (LLMs) have revolutionized natural language processing with their exceptional understanding, synthesizing, and reasoning capabilities. However, deploying LLMs on resource-constrained edge devices presents significant challenges due to computational limitations, memory constraints, and edge hardware heterogeneity. This survey provides a comprehensive overview of recent advancements in edge LLMs, covering the entire lifecycle: from resource-efficient model design and pre-deployment strategies to runtime inference optimizations. It also explores on-device applications across various domains. By synthesizing state-of-the-art techniques and identifying future research directions, this survey bridges the gap between the immense potential of LLMs and the constraints of edge computing.

Keywords

Cite

@article{arxiv.2410.11845,
  title  = {A Review on Edge Large Language Models: Design, Execution, and Applications},
  author = {Yue Zheng and Yuhao Chen and Bin Qian and Xiufang Shi and Yuanchao Shu and Jiming Chen},
  journal= {arXiv preprint arXiv:2410.11845},
  year   = {2025}
}

Comments

Accepted by ACM Computing Surveys(CSUR)

R2 v1 2026-06-28T19:23:00.475Z