English

A Systematic Characterization of LLM Inference on GPUs

Hardware Architecture 2025-12-02 v1

Abstract

This work presents a systematic characterization of Large Language Model (LLM) inference to address fragmented understanding. Through comprehensive experiments, we establish a four-dimensional analytical framework: (1) Two-Phase Heterogeneity Observation; (2) Microarchitectural Root Cause Analysis; (3) System Scaling Principles; and (4) Emerging Paradigm Boundaries. Our investigation progresses systematically from observation to foresight: identifying performance phenomena, revealing hardware causes, validating system behavior, and exploring new paradigms. This study not only consolidates a reliable empirical foundation for existing research but also provides new discoveries and practical optimization guidance for LLM inference.

Keywords

Cite

@article{arxiv.2512.01644,
  title  = {A Systematic Characterization of LLM Inference on GPUs},
  author = {Haonan Wang and Xuxin Xiao and Mingyu Yan and Zhuoyuan Zhu and Dengke Han and Duo Wang and Wenming Li and Xiaochun Ye and Cunchen Hu and Hongyang Chen and Guangyu Sun},
  journal= {arXiv preprint arXiv:2512.01644},
  year   = {2025}
}
R2 v1 2026-07-01T08:03:42.082Z