English

MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs

Computation and Language 2025-10-30 v2 Artificial Intelligence

Abstract

The widespread adoption of Large Language Models (LLMs) raises critical concerns about the factual accuracy of their outputs, especially in high-risk domains such as biomedicine, law, and education. Existing evaluation methods for short texts often fail on long-form content due to complex reasoning chains, intertwined perspectives, and cumulative information. To address this, we propose a systematic approach integrating large-scale long-form datasets, multi-agent verification mechanisms, and weighted evaluation metrics. We construct LongHalluQA, a Chinese long-form factuality dataset; and develop MAD-Fact, a debate-based multi-agent verification system. We introduce a fact importance hierarchy to capture the varying significance of claims in long-form texts. Experiments on two benchmarks show that larger LLMs generally maintain higher factual consistency, while domestic models excel on Chinese content. Our work provides a structured framework for evaluating and enhancing factual reliability in long-form LLM outputs, guiding their safe deployment in sensitive domains.

Keywords

Cite

@article{arxiv.2510.22967,
  title  = {MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs},
  author = {Yucheng Ning and Xixun Lin and Fang Fang and Yanan Cao},
  journal= {arXiv preprint arXiv:2510.22967},
  year   = {2025}
}

Comments

The article has been accepted by Frontiers of Computer Science (FCS), with the DOI: {10.1007/s11704-025-51369-x}

R2 v1 2026-07-01T07:07:03.183Z