English
Related papers

Related papers: Investigating differences in lab-quality and remot…

200 papers

We introduced a measurement procedure for the involuntary response of voice fundamental-frequency to frequency modulated auditory stimulation. This involuntary response plays an essential role in voice fundamental frequency control while…

A sound source was proposed for acoustic measurements of physical models of the human vocal tract. The physical models are produced by Fast Prototyping, based on Magnetic Resonance Imaging during prolonged vowel production. The sound…

Instrumentation and Detectors · Physics 2017-11-22 Antti Hannukainen , Juha Kuortti , Jarmo Malinen , Antti Ojalammi

Room acoustics analysis plays a central role in architectural design, audio engineering, speech intelligibility assessment, and hearing research. Despite the availability of standardized metrics such as reverberation time, clarity, and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-16 Mandip Goswami

Atrial fibrillation is increasingly prevalent, especially in the elderly, and challenging to detect due paroxysmal nature. Here, we propose novel computational methods based on heart beat intervals to facilitate rapid and robust…

Computers and Society · Computer Science 2017-11-30 Tamas Madl , David Madl

Self-supervised speech representations can hugely benefit downstream speech technologies, yet the properties that make them useful are still poorly understood. Two candidate properties related to the geometry of the representation space…

Computation and Language · Computer Science 2024-06-14 Mukhtar Mohamed , Oli Danyi Liu , Hao Tang , Sharon Goldwater

Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Hongyu Wang , Hui Li , Bo Li

We introduce and explore a new multimodal input representation for vision-language models: acoustic field video. Unlike conventional video (RGB with stereo/mono audio), our video stream provides a spatially grounded visualization of sound…

Human-Computer Interaction · Computer Science 2026-01-27 Daehwa Kim , Chris Harrison

While audio recordings in real life provide insights into social dynamics and conversational behavior, they also raise concerns about the privacy of personal, sensitive data. This article explores the effectiveness of restricting recordings…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-04 Jule Pohlhausen , Jörg Bitzer

Multimodal research and applications are becoming more commonplace as Virtual Reality (VR) technology integrates different sensory feedback, enabling the recreation of real spaces in an audio-visual context. Within VR experiences, numerous…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-08 Mauricio Flores-Vargas , Enda Bates , Rachel McDonnell

An approximation to coherent sampling, also known as boot-strapped waveform averaging, is presented. The method uses digital cavities to determine the condition for coherent sampling. It can be used to increase the effective sampling rate…

Data Analysis, Statistics and Probability · Physics 2018-04-04 Mattias Olsson , Fredrik Edman , Khadga Jung Karki

Audio-Visual Localization (AVL) aims to identify sound-emitting sources within a visual scene. However, existing studies focus on image-level audio-visual associations, failing to capture temporal dynamics. Moreover, they assume simplified…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Hahyeon Choi , Junhoo Lee , Nojun Kwak

Bioacoustic sensors, sometimes known as autonomous recording units (ARUs), can record sounds of wildlife over long periods of time in scalable and minimally invasive ways. Deriving per-species abundance estimates from these sensors requires…

The problem of audio-to-text alignment has seen significant amount of research using complete supervision during training. However, this is typically not in the context of long audio recordings wherein the text being queried does not appear…

Computation and Language · Computer Science 2023-10-11 Piyush Singh Pasi , Karthikeya Battepati , Preethi Jyothi , Ganesh Ramakrishnan , Tanmay Mahapatra , Manoj Singh

Objective: The Daily Phonotrauma Index (DPI) can quantify pathophysiological mechanisms associated with daily voice use in individuals with phonotraumatic vocal hyperfunction (PVH). Since DPI was developed based on week-long ambulatory…

Sound · Computer Science 2024-09-05 Hamzeh Ghasemzadeh , Robert E. Hillman , Jarrad H. Van Stan , Daryush D. Mehta

We show the theoretical and experimental combination of acoustic and optical methods for the in situ quantitative evaluation of the density, the viscosity and the thickness of soft layers adsorbed on chemically tailored metal surfaces. For…

Soft Condensed Matter · Physics 2007-05-23 L. Francis , J. -M. Friedt , C. Zhou , P. Bertrand

Processing long-form audio is a major challenge for Large Audio Language models (LALMs). These models struggle with the quadratic cost of attention ($O(N^2)$) and with modeling long-range temporal dependencies. Existing audio benchmarks are…

Masked speech modeling (MSM) methods such as wav2vec2 or w2v-BERT learn representations over speech frames which are randomly masked within an utterance. While these methods improve performance of Automatic Speech Recognition (ASR) systems,…

Audio-based classification techniques on body sounds have long been studied to aid in the diagnosis of respiratory diseases. While most research is centered on the use of cough as the main biomarker, other body sounds also have the…

Sound · Computer Science 2023-11-27 Tuan Truong , Matthias Lenga , Antoine Serrurier , Sadegh Mohammadi

Room impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional methods to estimate…

Sound · Computer Science 2022-11-28 Sagnik Majumder , Changan Chen , Ziad Al-Halah , Kristen Grauman

We address the problem of reconstructing articulatory movements, given audio and/or phonetic labels. The scarce availability of multi-speaker articulatory data makes it difficult to learn a reconstruction that generalizes to new speakers…

Computation and Language · Computer Science 2023-09-13 Rosanna Turrisi , Raffaele Tavarone , Leonardo Badino