English

Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks

Artificial Intelligence 2025-10-03 v2

Abstract

We study multi-turn multi-agent orchestration, where multiple large language model (LLM) agents interact over multiple turns by iteratively proposing answers or casting votes until reaching consensus. Using four LLMs (Gemini 2.5 Pro, GPT-5, Grok 4, and Claude Sonnet 4) on GPQA-Diamond, IFEval, and MuSR, we conduct two experiments: (i) benchmarking orchestration against single-LLM baselines; and (ii) ablations on GPQA-Diamond that vary whether agents see who authored answers and whether they can observe ongoing votes. Orchestration matches or exceeds the strongest single model and consistently outperforms the others. Analysis of best-achievable orchestration performance shows potential for further gains. The ablations show that revealing authorship increases self-voting and ties, and that showing ongoing votes amplifies herding, which speeds convergence but can sometimes yield premature consensus.

Keywords

Cite

@article{arxiv.2509.23537,
  title  = {Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks},
  author = {Aaron Xuxiang Tian and Ruofan Zhang and Jiayao Tang and Young Min Cho and Xueqian Li and Qiang Yi and Ji Wang and Zhunping Zhang and Danrui Qi and Zekun Li and Xingyu Xiang and Sharath Chandra Guntuku and Lyle Ungar and Tianyu Shi and Chi Wang},
  journal= {arXiv preprint arXiv:2509.23537},
  year   = {2025}
}

Comments

9 pages, 3 tables, 1 figure

R2 v1 2026-07-01T06:01:38.371Z