English

Open-Ended Multi-Modal Relational Reasoning for Video Question Answering

Artificial Intelligence 2024-06-12 v4 Human-Computer Interaction Robotics

Abstract

In this paper, we introduce a robotic agent specifically designed to analyze external environments and address participants' questions. The primary focus of this agent is to assist individuals using language-based interactions within video-based scenes. Our proposed method integrates video recognition technology and natural language processing models within the robotic agent. We investigate the crucial factors affecting human-robot interactions by examining pertinent issues arising between participants and robot agents. Methodologically, our experimental findings reveal a positive relationship between trust and interaction efficiency. Furthermore, our model demonstrates a 2\% to 3\% performance enhancement in comparison to other benchmark methods.

Keywords

Cite

@article{arxiv.2012.00822,
  title  = {Open-Ended Multi-Modal Relational Reasoning for Video Question Answering},
  author = {Haozheng Luo and Ruiyang Qin and Chenwei Xu and Guo Ye and Zening Luo},
  journal= {arXiv preprint arXiv:2012.00822},
  year   = {2024}
}

Comments

2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)

R2 v1 2026-06-23T20:39:15.430Z