English
Related papers

Related papers: Augmenting speech transcripts of VR recordings wit…

200 papers

Accurate point tracking in surgical environments remains challenging due to complex visual conditions, including smoke occlusion, specular reflections, and tissue deformation. While existing surgical tracking datasets provide coordinate…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Rulin Zhou , Wenlong He , An Wang , Jianhang Zhang , Xuanhui Zeng , Xi Zhang , Chaowei Zhu , Haijun Hu , Hongliang Ren

Robot teleoperation via augmented reality (AR) offers a promising path toward more intuitive human-robot interaction (HRI). We present a head-mounted AR 'puppeteer' system in which users control a physical robot by interacting with its…

Human-Computer Interaction · Computer Science 2026-05-19 Yuchong Zhang , Bastian Orthmann , Shichen Ji , Michael Welle , Jonne Van Haastregt , Danica Kragic

Visual storytelling is an emerging field that combines images and narratives to create engaging and contextually rich stories. Despite its potential, generating coherent and emotionally resonant visual stories remains challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Xiaochuan Lin , Xiangyong Chen

Multimodal automatic speech recognition systems integrate information from images to improve speech recognition quality, by grounding the speech in the visual context. While visual signals have been shown to be useful for recovering…

Computation and Language · Computer Science 2020-10-07 Tejas Srinivasan , Ramon Sanabria , Florian Metze , Desmond Elliott

Traditional sign language teaching methods face challenges such as limited feedback and diverse learning scenarios. Although 2D resources lack real-time feedback, classroom teaching is constrained by a scarcity of teacher. Methods based on…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Hongli Wen , Yang Xu , Lin Li , Xudong Ru , Xingce Wang , Zhongke Wu

Language models have been supervised with both language-only objective and visual grounding in existing studies of visual-grounded language learning. However, due to differences in the distribution and scale of visual-grounded datasets and…

Computation and Language · Computer Science 2024-01-10 Cong-Duy Nguyen , The-Anh Vu-Le , Thong Nguyen , Tho Quan , Luu Anh Tuan

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is…

Computer Vision and Pattern Recognition · Computer Science 2018-06-14 Aviv Gabbay , Asaph Shamir , Shmuel Peleg

We present RealityTalk, a system that augments real-time live presentations with speech-driven interactive virtual elements. Augmented presentations leverage embedded visuals and animation for engaging and expressive storytelling. However,…

Human-Computer Interaction · Computer Science 2022-08-15 Jian Liao , Adnan Karim , Shivesh Jadon , Rubaiat Habib Kazi , Ryo Suzuki

Data augmentation via voice conversion (VC) has been successfully applied to low-resource expressive text-to-speech (TTS) when only neutral data for the target speaker are available. Although the quality of VC is crucial for this approach,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-06 Ryo Terashima , Ryuichi Yamamoto , Eunwoo Song , Yuma Shirahata , Hyun-Wook Yoon , Jae-Min Kim , Kentaro Tachibana

The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image. However, some works suggest that introducing image information to model does not help…

Sound · Computer Science 2024-10-08 Jiliang Hu , Zuchao Li , Ping Wang , Haojun Ai , Lefei Zhang , Hai Zhao

Benchmarks for language-guided embodied agents typically assume text-based instructions, but deployed agents will encounter spoken instructions. While Automatic Speech Recognition (ASR) models can bridge the input gap, erroneous ASR…

Computation and Language · Computer Science 2023-10-11 Allen Chang , Xiaoyuan Zhu , Aarav Monga , Seoho Ahn , Tejas Srinivasan , Jesse Thomason

To explore a more scalable path for adding multimodal capabilities to existing LLMs, this paper addresses a fundamental question: Can a unimodal LLM, relying solely on text, reason about its own informational needs and provide effective…

Computation and Language · Computer Science 2026-01-13 Sazia Tabasum Mim , Jack Morris , Manish Dhakal , Yanming Xiu , Maria Gorlatova , Yi Ding

Anaphoric expressions, such as pronouns and referential descriptions, are situated with respect to the linguistic context of prior turns, as well as, the immediate visual environment. However, a speaker's referential descriptions do not…

Computation and Language · Computer Science 2023-07-27 Javier Chiyah-Garcia , Alessandro Suglia , José Lopes , Arash Eshghi , Helen Hastie

We present our work on the multimodal coreference resolution task of the Situated and Interactive Multimodal Conversation 2.0 (SIMMC 2.0) dataset as a part of the tenth Dialog System Technology Challenge (DSTC10). We propose a UNITER-based…

Computation and Language · Computer Science 2021-12-08 Yichen Huang , Yuchen Wang , Yik-Cheung Tam

Nonverbal communication is integral to human interaction, with gestures, facial expressions, and body language conveying critical aspects of intent and emotion. However, existing large language models (LLMs) fail to effectively incorporate…

Artificial Intelligence · Computer Science 2025-06-03 Youngmin Kim , Jiwan Chung , Jisoo Kim , Sunghyun Lee , Sangkyu Lee , Junhyeok Kim , Cheoljong Yang , Youngjae Yu

Virtual and augmented reality communication platforms are seen as promising modalities for next-generation remote face-to-face interactions. Our study attempts to explore non-verbal communication features in relation to their conversation…

Human-Computer Interaction · Computer Science 2019-04-10 Mohammad Keshavarzi , Michael Wu , Michael N. Chin , Robert N. Chin , Allen Y. Yang

Interactions with virtual assistants typically start with a predefined trigger phrase followed by the user command. To make interactions with the assistant more intuitive, we explore whether it is feasible to drop the requirement that users…

Computation and Language · Computer Science 2024-03-27 Dominik Wagner , Alexander Churchill , Siddharth Sigtia , Panayiotis Georgiou , Matt Mirsamadi , Aarshee Mishra , Erik Marchi

The dominant probing approaches rely on the zero-shot performance of image-text matching tasks to gain a finer-grained understanding of the representations learned by recent multimodal image-language transformer models. The evaluation is…

Computation and Language · Computer Science 2024-01-31 Ivana Beňová , Jana Košecká , Michal Gregor , Martin Tamajka , Marcel Veselý , Marián Šimko

Multimodal learning allows us to leverage information from multiple sources (visual, acoustic and text), similar to our experience of the real world. However, it is currently unclear to what extent auxiliary modalities improve performance…

Computation and Language · Computer Science 2020-01-01 Tejas Srinivasan , Ramon Sanabria , Florian Metze

Despite progress in Large Vision-Language Models (LVLMs), their capacity for visual reasoning is often limited by the binding problem: the failure to reliably associate perceptual features with their correct visual referents. This…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Amirmohammad Izadi , Mohammad Ali Banayeeanzade , Fatemeh Askari , Ali Rahimiakbar , Mohammad Mahdi Vahedi , Hosein Hasani , Mahdieh Soleymani Baghshah
‹ Prev 1 4 5 6 7 8 10 Next ›