English

Rethinking Artifact Evaluation for Software Engineering in the Age of Generative AI

Software Engineering 2026-04-21 v1

Abstract

Peer review in software engineering research operates under tight time constraints, while generative AI has substantially reduced the human effort required to produce polished research narratives. Reviewer attention is often spent on aspects of submissions such as writing quality or literature positioning that have become relatively less effort-intensive to address, rather than on evaluating the scientific substance of a paper. At the same time, assessing whether methods are implemented correctly, analyses are sound, and claims are supported by evidence remains effort-intensive and dependent on human expertise. In software engineering research, this substance is frequently embodied in artifacts, including code, data, evidence and analysis samples, and experimental infrastructure. In this position paper, we argue that artifact evaluation should be treated as a first-class component of peer review. We frame peer review as an attention allocation problem, examine how generative AI weakens narrative quality as a signal of rigor, and argue that artifact evaluation should play a more prominent role in peer review decisions.

Keywords

Cite

@article{arxiv.2604.16306,
  title  = {Rethinking Artifact Evaluation for Software Engineering in the Age of Generative AI},
  author = {Christoph Treude and Christopher M. Poskitt and Rashina Hoda},
  journal= {arXiv preprint arXiv:2604.16306},
  year   = {2026}
}

Comments

To appear in 2026 IEEE/ACM 48th International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE), April 12-18, 2026, Rio de Janeiro, Brazil