English

Ask No More: Deciding when to guess in referential visual dialogue

Computation and Language 2018-06-13 v2 Computer Vision and Pattern Recognition Multimedia

Abstract

Our goal is to explore how the abilities brought in by a dialogue manager can be included in end-to-end visually grounded conversational agents. We make initial steps towards this general goal by augmenting a task-oriented visual dialogue model with a decision-making component that decides whether to ask a follow-up question to identify a target referent in an image, or to stop the conversation to make a guess. Our analyses show that adding a decision making component produces dialogues that are less repetitive and that include fewer unnecessary questions, thus potentially leading to more efficient and less unnatural interactions.

Keywords

Cite

@article{arxiv.1805.06960,
  title  = {Ask No More: Deciding when to guess in referential visual dialogue},
  author = {Ravi Shekhar and Tim Baumgartner and Aashish Venkatesh and Elia Bruni and Raffaella Bernardi and Raquel Fernandez},
  journal= {arXiv preprint arXiv:1805.06960},
  year   = {2018}
}

Comments

COLING 2018 (accepted)

R2 v1 2026-06-23T01:59:16.563Z