English

Pointing out Human Answer Mistakes in a Goal-Oriented Visual Dialogue

Computer Vision and Pattern Recognition 2023-09-20 v1

Abstract

Effective communication between humans and intelligent agents has promising applications for solving complex problems. One such approach is visual dialogue, which leverages multimodal context to assist humans. However, real-world scenarios occasionally involve human mistakes, which can cause intelligent agents to fail. While most prior research assumes perfect answers from human interlocutors, we focus on a setting where the agent points out unintentional mistakes for the interlocutor to review, better reflecting real-world situations. In this paper, we show that human answer mistakes depend on question type and QA turn in the visual dialogue by analyzing a previously unused data collection of human mistakes. We demonstrate the effectiveness of those factors for the model's accuracy in a pointing-human-mistake task through experiments using a simple MLP model and a Visual Language Model.

Keywords

Cite

@article{arxiv.2309.10375,
  title  = {Pointing out Human Answer Mistakes in a Goal-Oriented Visual Dialogue},
  author = {Ryosuke Oshima and Seitaro Shinagawa and Hideki Tsunashima and Qi Feng and Shigeo Morishima},
  journal= {arXiv preprint arXiv:2309.10375},
  year   = {2023}
}

Comments

Accepted at ICCVW 2023

R2 v1 2026-06-28T12:25:45.780Z