English
Related papers

Related papers: Investigating differences in lab-quality and remot…

200 papers

We discuss post-processing of speech that has been recorded during Magnetic Resonance Imaging (MRI) of the vocal tract. Such speech recordings are contaminated by high levels of acoustic noise from the MRI scanner. Also, the frequency…

Sound · Computer Science 2016-06-22 Juha Kuortti , Jarmo Malinen , Antti Ojalammi

This article discusses aeroacoustic imaging methods based on correlation measurements in the frequency domain. Standard methods in this field assume that the estimated correlation matrix is superimposed with additive white noise. In this…

Signal Processing · Electrical Eng. & Systems 2020-12-30 Hans-Georg Raumer , Carsten Spehr , Thorsten Hohage , Daniel Ernst

While audio quality is a key performance metric for various audio processing tasks, including generative modeling, its objective measurement remains a challenge. Audio-Language Models (ALMs) are pre-trained on audio-text pairs that may…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-02 Soham Deshmukh , Dareen Alharthi , Benjamin Elizalde , Hannes Gamper , Mahmoud Al Ismail , Rita Singh , Bhiksha Raj , Huaming Wang

Audio quality assessment is critical for assessing the perceptual realism of sounds. However, the time and expense of obtaining ''gold standard'' human judgments limit the availability of such data. For AR&VR, good perceived sound quality…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-27 Pranay Manocha , Anurag Kumar , Buye Xu , Anjali Menon , Israel D. Gebru , Vamsi K. Ithapu , Paul Calamia

The objective of the present study is exploratory: to introduce and apply a new theory of speech rhythm zones or rhythm formants (R-formants). R-formants are zones of high magnitude frequencies in the low frequency (LF) long-term spectrum…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-13 Dafydd Gibbon , Peng Li

Since the COVID-19 pandemic in 2020, universities and companies have increasingly integrated hybrid features into their meeting spaces, or even created dedicated rooms for this purpose. While the importance of a fast and stable internet…

Computation and Language · Computer Science 2025-09-16 Robert Einig , Stefan Janscha , Jonas Schuster , Julian Koch , Martin Hagmueller , Barbara Schuppler

Audio-visual large language models (LLM) have drawn significant attention, yet the fine-grained combination of both input streams is rather under-explored, which is challenging but necessary for LLMs to understand general video inputs. To…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-11 Guangzhi Sun , Wenyi Yu , Changli Tang , Xianzhao Chen , Tian Tan , Wei Li , Lu Lu , Zejun Ma , Chao Zhang

Misdiagnosis can result in delayed treatment and patient harm. Robotic patient simulators (robopatients) provide a controlled framework for training and evaluating clinicians in rare and complex cases. We investigate auditory tactile…

ASR systems struggle with non-normative speech due to high acoustic variability and data scarcity. We propose a data-efficient method using phoneme-level uncertainty to guide fine-tuning for personalization. Instead of computationally…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Niclas Pokel , Pehuén Moure , Roman Böhringer , Yingqiang Gao

Passive acoustic monitoring can be an effective way of monitoring wildlife populations that are acoustically active but difficult to survey visually. Digital recorders allow surveyors to gather large volumes of data at low cost, but…

Sound · Computer Science 2023-08-25 Yuheng Wang , Juan Ye , David L. Borchers

Some optical measurements require relative timing of intensity variations with accuracy much finer than the camera frame period. One motivating example is dynamic aurora, where different prompt emissions are expected to originate from…

Space Physics · Physics 2026-05-29 Juha Vierinen , Pavithiran Sivasothy , Björn Gustavsson

Embedding acoustic information into fixed length representations is of interest for a whole range of applications in speech and audio technology. Two novel unsupervised approaches to generate acoustic embeddings by modelling of acoustic…

Computation and Language · Computer Science 2021-02-08 Yanpei Shi , Thomas Hain

The numerical investigation of acoustic damping materials, such as foams, constitutes a valuable enhancement to experimental testing. Typically, such materials are modeled in a homogenized way in order to reduce the computational effort and…

Numerical Analysis · Mathematics 2023-10-04 Lars Radtke , Paul Marter , Fabian Duvigneau , Sascha Eisenträger , Daniel Juhre , Alexander Düster

We performed 3D numerical simulations of the solar surface wave field for the quiet Sun and for three models with different localized sound-speed variations in the interior with: (i) deep, (ii) shallow, and (iii) two-layer structures. We…

Solar and Stellar Astrophysics · Physics 2012-09-24 Konstantin V. Parchevsky , Junwei Zhao , Thomas Hartlep , Alexander G. Kosovichev

Audiovisual synchronisation is the task of determining the time offset between speech audio and a video recording of the articulators. In child speech therapy, audio and ultrasound videos of the tongue are captured using instruments which…

Computation and Language · Computer Science 2019-11-28 Aciel Eshky , Manuel Sam Ribeiro , Korin Richmond , Steve Renals

Computer voice is experiencing a renaissance through the growing popularity of voice-based interfaces, agents, and environments. Yet, how to measure the user experience (UX) of voice-based systems remains an open and urgent question,…

Human-Computer Interaction · Computer Science 2021-03-15 Katie Seaborn , Jacqueline Urakami

This manuscript presents initial findings critical for supporting augmented acoustics experiments in custom-made hearing booths, addressing a key challenge in ensuring perceptual validity and experimental rigor in these highly sensitive…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Ernesto Accolti

Binaural recordings are a form of stereophonic recording method that replicates how human ears perceive sound, these types of recordings create a 3D aural image around the listener and are extremely immersive when well recorded and listened…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-29 Johann Kay Ann Tan

Background: The classroom discourse analysis has been transformed by the growing use of audio-video multimodal data, which demands analytical methods that balance interpretive depth with computational scalability. Methods: This study…

Physics and Society · Physics 2026-04-27 Vivek Upadhyay , Amaresh Chakrabarti

Automatic speaker verification (ASV) is the process to recognize persons using voice as biometric. The ASV systems show considerable recognition performance with sufficient amount of speech from matched condition. One of the crucial…

Multimedia · Computer Science 2018-12-04 Arnab Poddar , Md Sahidullah , Goutam Saha