AWS Bedrock 上的 LLM 性能分析:收据项目分类案例研究
人工智能
2026-04-03 v1 软件工程
摘要
本文提出了一套系统化、成本导向的大语言模型(LLM)评估框架,用于收据项目分类 within a production-oriented classification framework。我们比较了通过 AWS Bedrock 提供的四种指令调优模型:Claude 3.7 Sonnet、Claude 4 Sonnet、Mixtral 8x7B Instruct 和 Mistral 7B Instruct. 研究旨在(1)评估模型在准确率、响应稳定性和 token 级别成本方面的表现,(2)探讨在准确率和 incurred costs 两方面,零样本和 few-shot 提示方法哪种更合适。实验结果显示,Claude 3.7 Sonnet 在分类准确率和成本效率之间取得了最佳平衡。
关键词
引用
@article{arxiv.2604.01615,
title = {Analysis of LLM Performance on AWS Bedrock: Receipt-item Categorisation Case Study},
author = {Gabby Sanchez and Sneha Oommen and Cassandra T. Britto and Di Wang and Jung-De Chiou and Maria Spichkova},
journal= {arXiv preprint arXiv:2604.01615},
year = {2026}
}
备注
Preprint. Accepted to the 19th International Conference on Evaluation of Novel Approaches to Software Engineering (ENASE 2026). Final version to be published by SCITEPRESS, http://www.scitepress.org