增强二进制代码注释质量分类:集成生成式 AI 以提升准确率
软件工程
2023-10-19 v1 人工智能
机器学习
摘要
本报告聚焦于通过集成生成的代码与注释对来增强二进制代码注释质量分类模型,以提升模型准确率。数据集包含 9048 对用 C 编程语言编写的代码与注释,每对均被标注为“有用”或“无用”。此外,使用 Large Language Model Architecture(大型语言模型架构)生成代码与注释对,并对这些生成的对标注以表明其效用。此项工作的成果包含两种分类模型:一种使用原始数据集,另一种使用融合了新生成的代码注释对及其标签的增广数据集。
引用
@article{arxiv.2310.11467,
title = {Enhancing Binary Code Comment Quality Classification: Integrating Generative AI for Improved Accuracy},
author = {Rohith Arumugam S and Angel Deborah S},
journal= {arXiv preprint arXiv:2310.11467},
year = {2023}
}
备注
11 pages, 2 figures, 2 tables, Has been accepted for the Information Retrieval in Software Engineering track at Forum for Information Retrieval Evaluation 2023