English

Speed and Conversational Large Language Models: Not All Is About Tokens per Second

Computation and Language 2025-02-25 v1 Artificial Intelligence

Abstract

The speed of open-weights large language models (LLMs) and its dependency on the task at hand, when run on GPUs, is studied to present a comparative analysis of the speed of the most popular open LLMs.

Keywords

Cite

@article{arxiv.2502.16721,
  title  = {Speed and Conversational Large Language Models: Not All Is About Tokens per Second},
  author = {Javier Conde and Miguel González and Pedro Reviriego and Zhen Gao and Shanshan Liu and Fabrizio Lombardi},
  journal= {arXiv preprint arXiv:2502.16721},
  year   = {2025}
}
R2 v1 2026-06-28T21:54:48.009Z