English
Related papers

Related papers: Augmenting speech transcripts of VR recordings wit…

200 papers

In face-to-face interaction, we use multiple modalities, including speech and gestures, to communicate information and resolve references to objects. However, how representational co-speech gestures refer to objects remains understudied…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Esam Ghaleb , Bulat Khaertdinov , Aslı Özyürek , Raquel Fernández

Reference resolution, which aims to identify entities being referred to by a speaker, is more complex in real world settings: new referents may be created by processes the agents engage in and/or be salient only because they belong to the…

Computation and Language · Computer Science 2022-09-07 Abhinav Kumar , Barbara Di Eugenio , Abari Bhattacharya , Jillian Aurisano , Andrew Johnson

The visual dialog task requires an AI agent to interact with humans in multi-round dialogs based on a visual environment. As a common linguistic phenomenon, pronouns are often used in dialogs to improve the communication efficiency. As a…

Computation and Language · Computer Science 2022-05-31 Xintong Yu , Hongming Zhang , Ruixin Hong , Yangqiu Song , Changshui Zhang

We present a method for augmenting a Large Language Model (LLM) with a combination of text and visual data to enable accurate question answering in visualization of scientific data, making conversational visualization possible. LLMs…

Human-Computer Interaction · Computer Science 2025-01-17 Omar Mena , Alexandre Kouyoumdjian , Lonni Besançon , Michael Gleicher , Ivan Viola , Anders Ynnerman

Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated to the audio input. While previous approaches make crucial…

Sound · Computer Science 2022-04-29 Dan Oneata , Horia Cucu

Multimodal reference resolution, including phrase grounding, aims to understand the semantic relations between mentions and real-world objects. Phrase grounding between images and their captions is a well-established task. In contrast, for…

Computation and Language · Computer Science 2025-06-03 Shun Inadumi , Nobuhiro Ueda , Koichiro Yoshino

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

Visual dialog is a vision-language task where an agent needs to answer a series of questions grounded in an image based on the understanding of the dialog history and the image. The occurrences of coreference relations in the dialog makes…

Computer Vision and Pattern Recognition · Computer Science 2022-03-08 Mingxiao Li , Marie-Francine Moens

This paper introduces the concept of augmented conversation, which aims to support co-located in-person conversations via embedded speech-driven on-the-fly referencing in augmented reality (AR). Today computing technologies like smartphones…

Human-Computer Interaction · Computer Science 2024-05-30 Shivesh Jadon , Mehrad Faridan , Edward Mah , Rajan Vaish , Wesley Willett , Ryo Suzuki

Multimodal language models (MLLMs) are increasingly being applied in real-world environments, necessitating their ability to interpret 3D spaces and comprehend temporal dynamics. Current methods often rely on specialized architectural…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Benlin Liu , Yuhao Dong , Yiqin Wang , Zixian Ma , Yansong Tang , Luming Tang , Yongming Rao , Wei-Chiu Ma , Ranjay Krishna

Selection of occluded objects is a challenging problem in virtual reality, even more so if multiple objects are involved. With the advent of new artificial intelligence technologies, we explore the possibility of leveraging large language…

Human-Computer Interaction · Computer Science 2024-10-29 Junlong Chen , Jens Grubert , Per Ola Kristensson

Systems that can find correspondences between multiple modalities, such as between speech and images, have great potential to solve different recognition and data analysis tasks in an unsupervised manner. This work studies multimodal…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Khazar Khorrami , Okko Räsänen

3D visual grounding (3DVG), which aims to correlate a natural language description with the target object within a 3D scene, is a significant yet challenging task. Despite recent advancements in this domain, existing approaches commonly…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Xinyi Wang , Na Zhao , Zhiyuan Han , Dan Guo , Xun Yang

As augmented reality technology and hardware become more mature and affordable, researchers have been exploring more intuitive and discoverable interaction techniques for immersive environments. In this paper, we investigate multimodal…

Human-Computer Interaction · Computer Science 2022-12-02 Xiaoyan Zhou , Adam S. Williams , Francisco R. Ortega

In this paper, we present a novel dataset captured using a VR headset to record conversations between participants within a physics simulator (AI2-THOR). Our primary objective is to extend the field of co-speech gesture generation by…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Anna Deichler , Jim O'Regan , Jonas Beskow

Speech data has rich acoustic and paralinguistic information with important cues for understanding a speaker's tone, emotion, and intent, yet traditional large language models such as BERT do not incorporate this information. There has been…

Computation and Language · Computer Science 2023-11-14 Fatema Hasan , Yulong Li , James Foulds , Shimei Pan , Bishwaranjan Bhattacharjee

Vision-language models (VLMs) have shown to be effective at image retrieval based on simple text queries, but text-image retrieval based on conversational input remains a challenge. Consequently, if we want to use VLMs for reference…

Computation and Language · Computer Science 2023-09-26 Bram Willemsen , Livia Qian , Gabriel Skantze

Due to their significance in human communication, the automatic generation of co-speech gestures in artificial embodied agents has received a lot of attention. Although modern deep learning approaches can generate realistic-looking…

Human-Computer Interaction · Computer Science 2023-07-20 Hendric Voß , Stefan Kopp

Understanding 3D scenes from multi-view inputs has been proven to alleviate the view discrepancy issue in 3D visual grounding. However, existing methods normally neglect the view cues embedded in the text modality and fail to weigh the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Zoey Guo , Yiwen Tang , Ray Zhang , Dong Wang , Zhigang Wang , Bin Zhao , Xuelong Li

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in understanding multimodal inputs and have been widely integrated into Retrieval-Augmented Generation (RAG) based conversational systems. While current VLM-powered…

Computer Vision and Pattern Recognition · Computer Science 2025-01-23 Jingwei Yi , Junhao Yin , Ju Xu , Peng Bao , Yongliang Wang , Wei Fan , Hao Wang
‹ Prev 1 2 3 10 Next ›