中文

代码生成评估的基准测试与指标:批判性综述

人工智能 2024-06-19 v1 软件工程

摘要

随着大型语言模型(LLM)的快速发展,大量机器学习模型已被开发用于协助编程任务,包括从自然语言输入生成程序代码。然而,尽管已有大量研究致力于评估和比较这些工具,如何对LLMs进行评估以解决此类问题仍是未解之谜。本文对现有关于这些工具测试与评估工作进行了批判性综述,重点关注两个关键方面:用于评估的基准测试和指标。基于综述,进一步讨论了后续研究方向。

关键词

引用

@article{arxiv.2406.12655,
  title  = {Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review},
  author = {Debalina Ghosh Paul and Hong Zhu and Ian Bayley},
  journal= {arXiv preprint arXiv:2406.12655},
  year   = {2024}
}

备注

Accepted by the First IEEE International Workshop on Testing and Evaluation of Large Language Models (TELLMe 2024) and will be published in the proceedings of the IEEE AITest 2024 conference