English

Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond

Artificial Intelligence 2025-09-10 v3 Computation and Language Machine Learning

Abstract

The increasing demand for domain-specific evaluation of large language models (LLMs) has led to the development of numerous benchmarks. These efforts often adhere to the principle of data scaling, relying on large corpora or extensive question-answer (QA) sets to ensure broad coverage. However, the impact of corpus and QA set design on the precision and recall of domain-specific LLM performance remains poorly understood. In this paper, we argue that data scaling is not always the optimal principle for domain-specific benchmark construction. Instead, we introduce Comp-Comp, an iterative benchmarking framework grounded in the principle of comprehensiveness and compactness. Comprehensiveness ensures semantic recall by covering the full breadth of the domain, while compactness improves precision by reducing redundancy and noise. To demonstrate the effectiveness of our approach, we present a case study conducted at a well-renowned university, resulting in the creation of PolyBench, a large-scale, high-quality academic benchmark. Although this study focuses on academia, the Comp-Comp framework is domain-agnostic and readily adaptable to a wide range of specialized fields. The source code and datasets can be accessed at https://github.com/Anya-RB-Chen/COMP-COMP.

Keywords

Cite

@article{arxiv.2508.07353,
  title  = {Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond},
  author = {Rubing Chen and Jiaxin Wu and Jian Wang and Xulu Zhang and Wenqi Fan and Chenghua Lin and Xiao-Yong Wei and Qing Li},
  journal= {arXiv preprint arXiv:2508.07353},
  year   = {2025}
}

Comments

Accepted by EMNLP2025 Findings

R2 v1 2026-07-01T04:43:07.966Z