English
Related papers

Related papers: Voice Quality Dimensions as Interpretable Primitiv…

200 papers

Speech and voice conditions can alter the acoustic properties of speech, which could impact the performance of paralinguistic models for affect for people with atypical speech. We evaluate publicly available models for recognizing…

Machine Learning · Computer Science 2025-08-01 Jaya Narain , Amrit Romana , Vikramjit Mitra , Colin Lea , Shirley Ren

Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Matthieu Futeral , Andrea Agostinelli , Marco Tagliasacchi , Neil Zeghidour , Eugene Kharitonov

Subjective speech quality assessment is the gold standard for evaluating speech enhancement processing and telecommunication systems. The commonly used standard ITU-T Rec. P.800 defines how to measure speech quality in lab environments, and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-15 Babak Naderi , Ross Cutler , Nicolae-Catalin Ristea

Large audio and language models have recently demonstrated zero-shot reasoning capabilities across various domains. However, it remains unclear how the form of audio input, whether handcrafted acoustic features extracted from speech or the…

Sound · Computer Science 2026-05-26 Muhammad Ashad Kabir , Sirajam Munira

The quality of human voice plays an important role across various fields like music, speech therapy, and communication, yet it lacks a universally accepted, objective definition. Instead, voice quality is referred to using subjective…

Sound · Computer Science 2024-10-15 Hira Dhamyal , Rita Singh

Our voices are as distinctive as our faces and fingerprints. There is a spectrum of non-disjoint traits that make our voices unique and identifiable, such as the fundamental frequency, the intensity, and most interestingly the quality of…

Sound · Computer Science 2020-11-02 Shahan Ali Memon

With 4.5 million hours of English speech from 10 different sources across 120 countries and models of up to 10 billion parameters, we explore the frontiers of scale for automatic speech recognition. We propose data selection techniques to…

Computation and Language · Computer Science 2021-11-30 Alex Xiao , Weiyi Zheng , Gil Keren , Duc Le , Frank Zhang , Christian Fuegen , Ozlem Kalinli , Yatharth Saraf , Abdelrahman Mohamed

Diagnosing autism spectrum disorder (ASD) by identifying abnormal speech patterns from examiner-patient dialogues presents significant challenges due to the subtle and diverse manifestations of speech-related symptoms in affected…

Sound · Computer Science 2024-05-09 Chuanbo Hu , Jacob Thrasher , Wenqi Li , Mindi Ruan , Xiangxu Yu , Lynn K Paul , Shuo Wang , Xin Li

Recent advances in speech foundation models (SFMs) have enabled the direct processing of spoken language from raw audio, bypassing intermediate textual representations. This capability allows SFMs to be exposed to, and potentially respond…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-30 Harm Lameris , Shree Harsha Bokkahalli Satish , Joakim Gustafson , Éva Székely

Existing conversational datasets consist either of written proxies for dialog or small-scale transcriptions of natural speech. We introduce 'Interview': a large-scale (105K conversations) media dialog dataset collected from news interview…

Computation and Language · Computer Science 2020-04-08 Bodhisattwa Prasad Majumder , Shuyang Li , Jianmo Ni , Julian McAuley

Over the past few decades, computational methods have been developed to estimate perceptual audio quality. These methods, also referred to as objective quality measures, are usually developed and intended for a specific application domain.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-25 Matteo Torcoli , Thorsten Kastner , Jürgen Herre

This work presents self-supervised learning methods for developing monaural speaker-specific (i.e., personalized) speech enhancement models. While generalist models must broadly address many speakers, specialist models can adapt their…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-28 Aswin Sivaraman , Minje Kim

This study compares the performances of different algorithms for coding speech at low bit rates. In addition to widely deployed traditional vocoders, a selection of recently developed generative-model-based coders at different bit rates are…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-27 Wissam A. Jassim , Jan Skoglund , Michael Chinen , Andrew Hines

Prosodic cues in conversational speech aid listeners in discerning a message. We investigate whether acoustic cues in spoken dialogue can be used to identify the importance of individual words to the meaning of a conversation turn.…

Computation and Language · Computer Science 2019-07-18 Sushant Kafle , Cecilia O. Alm , Matt Huenerfauth

Diffusion models have found great success in generating high quality, natural samples of speech, but their potential for density estimation for speech has so far remained largely unexplored. In this work, we leverage an unconditional…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-16 Danilo de Oliveira , Julius Richter , Jean-Marie Lemercier , Simon Welker , Timo Gerkmann

Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-28 Wenming Tu , Guanrou Yang , Ruiqi Yan , Wenxi Chen , Ziyang Ma , Yipeng Kang , Kai Yu , Xie Chen , Zilong Zheng

Speaker embeddings are widely used in speaker verification systems and other applications where it is useful to characterise the voice of a speaker with a fixed-length vector. These embeddings tend to be treated as "black box" encodings,…

Sound · Computer Science 2025-10-21 Mark Huckvale

Recent speech enhancement models have shown impressive performance gains by scaling up model complexity and training data. However, the impact of dataset variability (e.g. text, language, speaker, and noise) has been underexplored.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-20 Leying Zhang , Wangyou Zhang , Chenda Li , Yanmin Qian

We introduce a new unsupervised task, spoken language modeling: the learning of linguistic representations from raw audio signals without any labels, along with the Zero Resource Speech Benchmark 2021: a suite of 4 black-box, zero-shot…

Computation and Language · Computer Science 2020-12-02 Tu Anh Nguyen , Maureen de Seyssel , Patricia Rozé , Morgane Rivière , Evgeny Kharitonov , Alexei Baevski , Ewan Dunbar , Emmanuel Dupoux

Modelling of early language acquisition aims to understand how infants bootstrap their language skills. The modelling encompasses properties of the input data used for training the models, the cognitive hypotheses and their algorithmic…

Computation and Language · Computer Science 2023-05-04 María Andrea Cruz Blandón , Alejandrina Cristia , Okko Räsänen
‹ Prev 1 2 3 10 Next ›