English

Argument Summarization and its Evaluation in the Era of Large Language Models

Computation and Language 2025-10-10 v4

Abstract

Large Language Models (LLMs) have revolutionized various Natural Language Generation (NLG) tasks, including Argument Summarization (ArgSum), a key subfield of Argument Mining. This paper investigates the integration of state-of-the-art LLMs into ArgSum systems and their evaluation. In particular, we propose a novel prompt-based evaluation scheme, and validate it through a novel human benchmark dataset. Our work makes three main contributions: (i) the integration of LLMs into existing ArgSum systems, (ii) the development of two new LLM-based ArgSum systems, benchmarked against prior methods, and (iii) the introduction of an advanced LLM-based evaluation scheme. We demonstrate that the use of LLMs substantially improves both the generation and evaluation of argument summaries, achieving state-of-the-art results and advancing the field of ArgSum. We also show that among the four LLMs integrated in (i) and (ii), Qwen-3-32B, despite having the fewest parameters, performs best, even surpassing GPT-4o.

Keywords

Cite

@article{arxiv.2503.00847,
  title  = {Argument Summarization and its Evaluation in the Era of Large Language Models},
  author = {Moritz Altemeyer and Steffen Eger and Johannes Daxenberger and Yanran Chen and Tim Altendorf and Philipp Cimiano and Benjamin Schiller},
  journal= {arXiv preprint arXiv:2503.00847},
  year   = {2025}
}

Comments

EMNLP 2025 Main Camera-ready