代码生成评估的基准测试与指标:批判性综述
人工智能
2024-06-19 v1 软件工程
摘要
随着大型语言模型(LLM)的快速发展,大量机器学习模型已被开发用于协助编程任务,包括从自然语言输入生成程序代码。然而,尽管已有大量研究致力于评估和比较这些工具,如何对LLMs进行评估以解决此类问题仍是未解之谜。本文对现有关于这些工具测试与评估工作进行了批判性综述,重点关注两个关键方面:用于评估的基准测试和指标。基于综述,进一步讨论了后续研究方向。
引用
@article{arxiv.2406.12655,
title = {Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review},
author = {Debalina Ghosh Paul and Hong Zhu and Ian Bayley},
journal= {arXiv preprint arXiv:2406.12655},
year = {2024}
}
备注
Accepted by the First IEEE International Workshop on Testing and Evaluation of Large Language Models (TELLMe 2024) and will be published in the proceedings of the IEEE AITest 2024 conference