English

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

Multiagent Systems 2026-08-04 v1 Computation and Language

Abstract

Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is lost when user requests are converted into executable plans. We study the planner in a multi-agent system as the request-to-action interface and derive an actionable taxonomy of planning-grounding failures from failed real-world task executions. LLM-based analysis shows that these failures constitute an increasing share of unsuccessful executions as language-resource availability declines, with the strongest effects in low-resource languages. To test whether the taxonomy supports mitigation, we introduce TART, Taxonomy-Guided Actionable Representation, that makes the taxonomy's key aspects explicit to the planner and downstream sub-agents. Across multiple languages, three LLM backbones, two datasets, and two agentic configurations, TART consistently improves performance. On multilingual GAIA, it raises a state-of-the-art system's accuracy by 5.6 percentage points averaged across eleven languages spanning low- to high-resource settings.

Keywords

Cite

@article{arxiv.2608.03735,
  title  = {An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures},
  author = {Vikas Pahuja and Jonathan Brokman and Omer Hofman and Tamir Nizri and Daniel Vishna and Seraphina Goldfarb-Tarrant and Kelly Marchisio and Hisashi Kojima and Roman Vainshtein},
  journal= {arXiv preprint arXiv:2608.03735},
  year   = {2026}
}

Comments

22 pages, 11 figures