English
Related papers

Related papers: Investigating differences in lab-quality and remot…

200 papers

We propose a cloud-based multimodal dialog platform for the remote assessment and monitoring of Amyotrophic Lateral Sclerosis (ALS) at scale. This paper presents our vision, technology setup, and an initial investigation of the efficacy of…

Distance Metric Learning (DML) has typically dominated the audio-visual speaker verification problem space, owing to strong performance in new and unseen classes. In our work, we explored multitask learning techniques to further enhance…

Sound · Computer Science 2024-09-25 Anith Selvakumar , Homa Fashandi

Most mainstream Automatic Speech Recognition (ASR) systems consider all feature frames equally important. However, acoustic landmark theory is based on a contradictory idea, that some frames are more important than others. Acoustic landmark…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-04 Di He , Boon Pang Lim , Xuesong Yang , Mark Hasegawa-Johnson , Deming Chen

Atomic force microscope (AFM) users often calibrate the spring constants of cantilevers using functionality built into individual instruments. This is performed without reference to a global standard, which hinders robust comparison of…

Large language models often suffer from fact loss, timeline confusion, persona drift, and reduced stability during long-range interaction, especially under high-noise knowledge bases, context clearing, and cross-model transfer. To address…

Artificial Intelligence · Computer Science 2026-05-15 Zhao Yang , Wang Huan , Li Yingshuo , Tu Haomiao , Lin Hujite

Measuring room impulse responses (RIRs) at multiple spatial points is a time-consuming task, while simulations require detailed knowledge of the room's acoustic environment. In prior work, we proposed a method for estimating the early part…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-16 Kathleen MacWilliam , Thomas Dietzen , Toon van Waterschoot

Audio LLMs have shown a strong ability to understand audio samples, yet their reliability in complex acoustic scenes remains under-explored. Unlike prior work limited to small scale or less controlled query construction, we present a…

Sound · Computer Science 2026-03-05 Taehan Lee , Jaehan Jung , Hyukjun Lee

We present a phoneme-level analysis of automatic speech recognition (ASR) for two low-resourced and phonologically complex East Caucasian languages, Archi and Rutul, based on curated and standardized speech-transcript resources totaling…

Computation and Language · Computer Science 2026-04-21 V. S. D. S. Mahesh Akavarapu , Michael Daniel , Gerhard Jäger

Voice activity detection (VAD) is a challenging task in low signal-to-noise ratio (SNR) environment, especially in non-stationary noise. To deal with this issue, we propose a novel attention module that can be integrated in Long Short-Term…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-26 Joohyung Lee , Youngmoon Jung , Hoirin Kim

Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propose the use of natural…

Sound · Computer Science 2023-12-08 Hira Dhamyal , Benjamin Elizalde , Soham Deshmukh , Huaming Wang , Bhiksha Raj , Rita Singh

Due to the high demand for mobile applications, given the exponential growth of users of this type of technology, testing professionals are frequently required to invest time in studying testing tools, in particular, because nowadays,…

Software Engineering · Computer Science 2023-07-04 Gustavo da Silva , Ronnie de Souza Santos

The historical and geographical spread from older to more modern languages has long been studied by examining textual changes and in terms of changes in phonetic transcriptions. However, it is more difficult to analyze language change from…

Applications · Statistics 2017-05-19 Davide Pigoli , Pantelis Z. Hadjipantelis , John S. Coleman , John A. D. Aston

In clinical voice signal analysis, mishandling of subharmonic voicing may cause an acoustic parameter to signal false negatives. As such, the ability of a fundamental frequency estimator to identify speaking fundamental frequency is…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-10 Takeshi Ikuma , Melda Kunduk , Andrew J. McWhorter

The accessibility of long-duration recorders, adapted to sometimes demanding field conditions, has enabled the deployment of extensive animal population monitoring campaigns through ecoacoustics. The effectiveness of automatic signal…

Sound · Computer Science 2025-07-04 Jérémy Rouch , M Ducrettet , S Haupert , R Emonet , F Sèbe

In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR)…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Xiulong Liu , Anurag Kumar , Paul Calamia , Sebastia V. Amengual , Calvin Murdock , Ishwarya Ananthabhotla , Philip Robinson , Eli Shlizerman , Vamsi Krishna Ithapu , Ruohan Gao

The Frequency Following Response (FFR) reflects the brain's neural encoding of auditory stimuli including speech. Because the fundamental frequency (F0), a physical correlate of pitch, is one of the essential features of speech, there has…

We investigate the acoustical properties of uncompressed and compressed open-celled aluminum metal foams fabricated using a directional solidification foaming process. We compressed the fabricated foams using a hydraulic press to different…

Applied Physics · Physics 2021-11-08 Amulya Lomte , Bhisham Sharma , Mary Drouin , Denver Schaffarzick

When it comes to authentication in speaker verification systems, not all utterances are created equal. It is essential to estimate the quality of test utterances in order to account for varying acoustic conditions. In addition to the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-12 Nicholas Klein , Ganesh Sivaraman , Elie Khoury

Passive acoustic mapping enables the spatial mapping and temporal monitoring of cavitation activity, playing a crucial role in therapeutic ultrasound applications. Most conventional beamforming methods, whether implemented in the time or…

Signal Processing · Electrical Eng. & Systems 2025-11-26 Tatiana Gelvez-Barrera , Barbara Nicolas , Denis Kouamé , Bruno Gilles , Adrian Basarab

While automatic speech recognition (ASR) greatly benefits from data augmentation, the augmentation recipes themselves tend to be heuristic. In this paper, we address one of the heuristic approach associated with balancing the right amount…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-17 Vishwanath Pratap Singh , Federico Malato , Ville Hautamaki , Md. Sahidullah , Tomi Kinnunen