English

An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue

Computation and Language 2025-03-19 v2 Artificial Intelligence Sound Audio and Speech Processing

Abstract

Handling multi-party dialogues represents a significant step for advancing spoken dialogue systems, necessitating the development of tasks specific to multi-party interactions. To address this challenge, we are constructing a multi-modal multi-party dialogue corpus of triadic (three-participant) discussions. This paper focuses on the task of addressee recognition, identifying who is being addressed to take the next turn, a critical component unique to multi-party dialogue systems. A subset of the corpus was annotated with addressee information, revealing that explicit addressees are indicated in approximately 20% of conversational turns. To evaluate the task's complexity, we benchmarked the performance of a large language model (GPT-4o) on addressee recognition. The results showed that GPT-4o achieved an accuracy only marginally above chance, underscoring the challenges of addressee recognition in multi-party dialogue. These findings highlight the need for further research to enhance the capabilities of large language models in understanding and navigating the intricacies of multi-party conversational dynamics.

Keywords

Cite

@article{arxiv.2501.16643,
  title  = {An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue},
  author = {Koji Inoue and Divesh Lala and Mikey Elmers and Keiko Ochi and Tatsuya Kawahara},
  journal= {arXiv preprint arXiv:2501.16643},
  year   = {2025}
}

Comments

This paper has been accepted for presentation at International Workshop on Spoken Dialogue Systems Technology 2025 (IWSDS 2025) and represents the author's version of the work

R2 v1 2026-06-28T21:21:09.982Z