English

A study of the impact of generative AI-based data augmentation on software metadata classification

Software Engineering 2023-10-24 v1 Artificial Intelligence Computation and Language Machine Learning

Abstract

This paper presents the system submitted by the team from IIT(ISM) Dhanbad in FIRE IRSE 2023 shared task 1 on the automatic usefulness prediction of code-comment pairs as well as the impact of Large Language Model(LLM) generated data on original base data towards an associated source code. We have developed a framework where we train a machine learning-based model using the neural contextual representations of the comments and their corresponding codes to predict the usefulness of code-comments pair and performance analysis with LLM-generated data with base data. In the official assessment, our system achieves a 4% increase in F1-score from baseline and the quality of generated data.

Keywords

Cite

@article{arxiv.2310.13714,
  title  = {A study of the impact of generative AI-based data augmentation on software metadata classification},
  author = {Tripti Kumari and Chakali Sai Charan and Ayan Das},
  journal= {arXiv preprint arXiv:2310.13714},
  year   = {2023}
}
R2 v1 2026-06-28T12:57:11.209Z