中文

速度与对话型大语言模型:并非所有问题都关于每秒标记数

计算与语言 2025-02-25 v1 人工智能

摘要

本文研究开源大语言模型(LLM)的速度及其在 GPU 上运行时的任务依赖性,旨在对最流行的开源 LLM 速度进行比较分析。

关键词

引用

@article{arxiv.2502.16721,
  title  = {Speed and Conversational Large Language Models: Not All Is About Tokens per Second},
  author = {Javier Conde and Miguel González and Pedro Reviriego and Zhen Gao and Shanshan Liu and Fabrizio Lombardi},
  journal= {arXiv preprint arXiv:2502.16721},
  year   = {2025}
}