English

Views Are My Own, but Also Yours: Benchmarking Theory of Mind Using Common Ground

Computation and Language 2024-06-11 v2

Abstract

Evaluating the theory of mind (ToM) capabilities of language models (LMs) has recently received a great deal of attention. However, many existing benchmarks rely on synthetic data, which risks misaligning the resulting experiments with human behavior. We introduce the first ToM dataset based on naturally occurring spoken dialogs, Common-ToM, and show that LMs struggle to demonstrate ToM. We then show that integrating a simple, explicit representation of beliefs improves LM performance on Common-ToM.

Keywords

Cite

@article{arxiv.2403.02451,
  title  = {Views Are My Own, but Also Yours: Benchmarking Theory of Mind Using Common Ground},
  author = {Adil Soubki and John Murzaku and Arash Yousefi Jordehi and Peter Zeng and Magdalena Markowska and Seyed Abolghasem Mirroshandel and Owen Rambow},
  journal= {arXiv preprint arXiv:2403.02451},
  year   = {2024}
}