English

Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7

Computation and Language 2018-06-05 v1 Computer Vision and Pattern Recognition

Abstract

Scene-aware dialog systems will be able to have conversations with users about the objects and events around them. Progress on such systems can be made by integrating state-of-the-art technologies from multiple research areas including end-to-end dialog systems visual dialog, and video description. We introduce the Audio Visual Scene Aware Dialog (AVSD) challenge and dataset. In this challenge, which is one track of the 7th Dialog System Technology Challenges (DSTC7) workshop1, the task is to build a system that generates responses in a dialog about an input video

Keywords

Cite

@article{arxiv.1806.00525,
  title  = {Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7},
  author = {Huda Alamri and Vincent Cartillier and Raphael Gontijo Lopes and Abhishek Das and Jue Wang and Irfan Essa and Dhruv Batra and Devi Parikh and Anoop Cherian and Tim K. Marks and Chiori Hori},
  journal= {arXiv preprint arXiv:1806.00525},
  year   = {2018}
}
R2 v1 2026-06-23T02:16:38.251Z