English
Related papers

Related papers: Audio-Visual Scene-Aware Dialog

200 papers

Voice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a…

Sound · Computer Science 2021-05-11 Heinrich Dinkel , Shuai Wang , Xuenan Xu , Mengyue Wu , Kai Yu

In this work, following the intuition that adverbs describing scene-sequences are best identified by reasoning over high-level concepts of object-behavior, we propose the design of a new framework that reasons over object-behaviours…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Amrit Diggavi Seshadri , Alessandra Russo

Most current AI systems rely on the premise that the input visual data are sufficient to achieve competitive performance in various computer vision tasks. However, the classic task setup rarely considers the challenging, yet common…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Zhenghao Zhao , Ye Zhu , Xiaoguang Zhu , Yuzhang Shang , Yan Yan

Visual dialog is a task of answering a series of inter-dependent questions given an input image, and often requires to resolve visual references among the questions. This problem is different from visual question answering (VQA), which…

Computer Vision and Pattern Recognition · Computer Science 2018-08-08 Paul Hongsuck Seo , Andreas Lehrmann , Bohyung Han , Leonid Sigal

Vision-language models have shown impressive progress in recent years. However, existing models are largely limited to turn-based interactions, where each turn must be stepped (i.e., prompted) by the user. Open-ended, asynchronous…

In human conversation, empathic dialogue requires nuanced temporal cues indicating whether the conversational partner is paying attention. This type of "active listening" is overlooked in the design of Conversational Agents (CAs), which use…

Human-Computer Interaction · Computer Science 2026-02-09 Zhihan Jiang , Qianhui Chen , Chu Zhang , Yanheng Li , Ray LC

We present a new research task and a dataset to understand human social interactions via computational methods, to ultimately endow machines with the ability to encode and decode a broad channel of social signals humans use. This research…

Computer Vision and Pattern Recognition · Computer Science 2019-06-11 Hanbyul Joo , Tomas Simon , Mina Cikara , Yaser Sheikh

Given a video with aligned dialogue, people can often infer what is more likely to happen next. Making such predictions requires not only a deep understanding of the rich dynamics underlying the video and dialogue, but also a significant…

Computation and Language · Computer Science 2020-10-19 Jie Lei , Licheng Yu , Tamara L. Berg , Mohit Bansal

As audio-visual systems increasingly bring immersive and interactive capabilities into our work and leisure activities, so the need for naturalistic test material grows. New volumetric datasets have captured high-quality 3D video, but…

Multimedia · Computer Science 2021-05-04 Hanne Stenzel , Davide Berghi , Marco Volino , Philip J. B. Jackson

The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have explored talking or…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Joon Son Chung , Shinji Watanabe

People interacting with voice assistants are often frustrated by voice assistants' frequent errors and inability to respond to backchannel cues. We introduce an open-source video dataset of 21 participants' interactions with a voice…

Human-Computer Interaction · Computer Science 2021-04-16 Andrea Cuadra , Hansol Lee , Jason Cho , Wendy Ju

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Daniel Cudeiro , Timo Bolkart , Cassidy Laidlaw , Anurag Ranjan , Michael J. Black

Recent advances in conversational AI have been substantial, but developing real-time systems for perceptual task guidance remains challenging. These systems must provide interactive, proactive assistance based on streaming visual inputs,…

Artificial Intelligence · Computer Science 2025-06-09 Yichi Zhang , Xin Luna Dong , Zhaojiang Lin , Andrea Madotto , Anuj Kumar , Babak Damavandi , Joyce Chai , Seungwhan Moon

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

We present a novel approach to multilingual audio-visual speech recognition tasks by introducing a single model on a multilingual dataset. Motivated by a human cognitive system where humans can intuitively distinguish different languages…

Multimedia · Computer Science 2023-10-24 Joanna Hong , Se Jin Park , Yong Man Ro

We introduce a video framework for modeling the association between verbal and non-verbal communication during dyadic conversation. Given the input speech of a speaker, our approach retrieves a video of a listener, who has facial…

Computer Vision and Pattern Recognition · Computer Science 2023-01-27 Scott Geng , Revant Teotia , Purva Tendulkar , Sachit Menon , Carl Vondrick

This paper presents a dataset collected from natural dialogs which enables to test the ability of dialog systems to learn new facts from user utterances throughout the dialog. This interactive learning will help with one of the most…

Computation and Language · Computer Science 2016-05-17 Miroslav Vodolán , Filip Jurčíček

In order to better simulate the real human conversation process, models need to generate dialogue utterances based on not only preceding textual contexts but also visual contexts. However, with the development of multi-modal dialogue…

Computation and Language · Computer Science 2021-09-29 Shuhe Wang , Yuxian Meng , Xiaoya Li , Xiaofei Sun , Rongbin Ouyang , Jiwei Li

A goal-oriented visual dialogue involves multi-turn interactions between two agents, Questioner and Oracle. During which, the answer given by Oracle is of great significance, as it provides golden response to what Questioner concerns. Based…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Zipeng Xu , Fangxiang Feng , Xiaojie Wang , Yushu Yang , Huixing Jiang , Zhongyuan Wang

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Ruohan Gao , Kristen Grauman