English
Related papers

Related papers: A virtual reality-based method for examining audio…

200 papers

The audio visual benefit in speech perception, where congruent visual input enhances auditory processing, is well documented across age groups, particularly in challenging listening conditions and among individuals with varying hearing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-15 Mareike Daeglau , Juergen Otten , Giso Grimm , Bojana Mirkovic , Volker Hohmann , Stefan Debener

Recognizing who is speaking in a crowded scene is a key challenge towards the understanding of the social interactions going on within. Detecting speaking status from body movement alone opens the door for the analysis of social scenes in…

Computer Vision and Pattern Recognition · Computer Science 2022-11-02 Jose Vargas-Quiros , Laura Cabrera-Quiros , Hayley Hung

Counseling is carried out as spoken conversation between a therapist and a client. The empathy level expressed by the therapist is considered an important index of the quality of counseling and often assessed by an observer or the client.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-24 Dehua Tao , Tan Lee , Harold Chui , Sarah Luk

Many people with some form of hearing loss consider lipreading as their primary mode of day-to-day communication. However, finding resources to learn or improve one's lipreading skills can be challenging. This is further exacerbated in the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Aditya Agarwal , Bipasha Sen , Rudrabha Mukhopadhyay , Vinay Namboodiri , C. V Jawahar

We address the problem of inferring a speaker's level of certainty based on prosodic information in the speech signal, which has application in speech-based dialogue systems. We show that using phrase-level prosodic features centered around…

Computation and Language · Computer Science 2011-03-11 Heather Pon-Barry , Stuart M. Shieber

In order to create effective storytelling agents three fundamental questions must be answered: first, is a physically embodied agent preferable to a virtual agent or a voice-only narration? Second, does a human voice have an advantage over…

Robotics · Computer Science 2016-07-20 Sandra Costa , Alberto Brunete , Byung-Chull Bae , Nikolaos Mavridis

With the widespread use of intelligent systems, such as smart speakers, addressee recognition has become a concern in human-computer interaction, as more and more people expect such systems to understand complicated social scenes, including…

Artificial Intelligence · Computer Science 2018-09-13 Thao Minh Le , Nobuyuki Shimizu , Takashi Miyazaki , Koichi Shinoda

Haptic feedback is an important component of creating an immersive mixed reality experience. Traditionally, haptic forces are rendered in response to the user's interactions with the virtual environment. In this work, we explore the idea of…

Robotics · Computer Science 2023-03-01 Niall L. Williams , Nicholas Rewkowski , Jiasheng Li , Ming C. Lin

Cochlear implants (CIs) are devices that restore the sense of hearing in people with severe sensorineural hearing loss. An electrode array inserted in the cochlea bypasses the natural transducer mechanism that transforms mechanical sound…

Neurons and Cognition · Quantitative Biology 2025-01-30 Franklin Alvarez , Yixuan Zhang , Daniel Kipping , Waldo Nogueira

Access to non-verbal cues in social interactions is vital for people with visual impairment. It has been shown that non-verbal cues such as eye contact, number of people, their names and positions are helpful for individuals who are blind.…

Computers and Society · Computer Science 2017-11-30 M. Saquib Sarfraz , Angela Constantinescu , Melanie Zuzej , Rainer Stiefelhagen

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-21 Abhinav Shukla , Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

Virtual reality labs for hearing research are commonly designed to achieve maximal acoustical accuracy of virtual environments. For a high immersion, 3D video systems are applied, that ideally do not influence the acoustical conditions. In…

Medical Physics · Physics 2023-09-22 Jan Heeren , Giso Grimm , Stephan Ewert , Volker Hohmann

Parsing spoken dialogue poses unique difficulties, including disfluencies and unmarked boundaries between sentence-like units. Previous work has shown that prosody can help with parsing disfluent speech (Tran et al. 2018), but has assumed…

Computation and Language · Computer Science 2021-10-13 Elizabeth Nielsen , Mark Steedman , Sharon Goldwater

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to…

Computer Vision and Pattern Recognition · Computer Science 2022-06-15 Changan Chen , Ruohan Gao , Paul Calamia , Kristen Grauman

Visual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The…

Sound · Computer Science 2023-02-23 Meng Liu , Kong Aik Lee , Longbiao Wang , Hanyi Zhang , Chang Zeng , Jianwu Dang

The task of classifying emotions within a musical track has received widespread attention within the Music Information Retrieval (MIR) community. Music emotion recognition has traditionally relied on the use of acoustic features, verbal…

Sound · Computer Science 2021-06-15 Nicholas Farris , Brian Model , Richard Savery , Gil Weinberg

Existing robotic manipulation methods primarily rely on visual and proprioceptive observations, which may struggle to infer contact-related interaction states in partially observable real-world environments. Acoustic cues, by contrast,…

Robotics · Computer Science 2026-02-17 Siyuan Li , Jiani Lu , Yu Song , Xianren Li , Bo An , Peng Liu

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

The labelling of speech corpora is a laborious and time-consuming process. The ProsoBeast Annotation Tool seeks to ease and accelerate this process by providing an interactive 2D representation of the prosodic landscape of the data, in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Branislav Gerazov , Michael Wagner

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current…

Multimedia · Computer Science 2025-05-29 Yong Ren , Chenxing Li , Le Xu , Hao Gu , Duzhen Zhang , Yujie Chen , Manjie Xu , Ruibo Fu , Shan Yang , Dong Yu