English
Related papers

Related papers: Deep Learning-based Non-Intrusive Multi-Objective …

200 papers

The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in the mean opinion…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-06 Meng Yu , Chunlei Zhang , Yong Xu , Shixiong Zhang , Dong Yu

Deep neural network based speech enhancement technique focuses on learning a noisy-to-clean transformation supervised by paired training data. However, the task-specific evaluation metric (e.g., PESQ) is usually non-differentiable and can…

Sound · Computer Science 2023-02-24 Chen Chen , Yuchen Hu , Weiwei Weng , Eng Siong Chng

Data-driven speech enhancement employing deep neural networks (DNNs) can provide state-of-the-art performance even in the presence of non-stationary noise. During the training process, most of the speech enhancement neural networks are…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-01 Ziyi Xu , Maximilian Strake , Tim Fingscheidt

The majority of deep neural network (DNN) based speech enhancement algorithms rely on the mean-square error (MSE) criterion of short-time spectral amplitudes (STSA), which has no apparent link to human perception, e.g. speech…

Sound · Computer Science 2018-12-05 Morten Kolbæk , Zheng-Hua Tan , Jesper Jensen

Silent Speech Interfaces (SSIs) offer a noninvasive alternative to brain-computer interfaces for soundless verbal communication. We introduce Multimodal Orofacial Neural Audio (MONA), a system that leverages cross-modal alignment through…

Human-Computer Interaction · Computer Science 2024-03-12 Tyler Benster , Guy Wilson , Reshef Elisha , Francis R Willett , Shaul Druckmann

Speech quality assessment has been a critical component in many voice communication related applications such as telephony and online conferencing. Traditional intrusive speech quality assessment requires the clean reference of the degraded…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-07 Yuchen Liu , Li-Chia Yang , Alex Pawlicki , Marko Stamenovic

Speech quality assessment (SQA) aims to evaluate the quality of speech samples without relying on time-consuming listener questionnaires. Recent efforts have focused on training neural-based SQA models to predict the mean opinion score…

Sound · Computer Science 2025-06-24 Yuto Kondo , Hirokazu Kameoka , Kou Tanaka , Takuhiro Kaneko

The acoustic environment can degrade speech quality during communication (e.g., video call, remote presentation, outside voice recording), and its impact is often unknown. Objective metrics for speech quality have proven challenging to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-30 Karl El Hajal , Milos Cernak , Pablo Mainar

The mean opinion score (MOS) is a standard metric for assessing speech quality, but its singular focus fails to identify specific distortions when low scores are observed. The NISQA dataset addresses this limitation by providing ratings…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-06 Fredrik Cumlin , Xinyu Liang , Victor Ungureanu , Chandan K. A. Reddy , Christian Schüldt , Saikat Chatterjee

Speech synthesis quality prediction has made remarkable progress with the development of supervised and self-supervised learning (SSL) MOS predictors but some aspects related to the data are still unclear and require further study. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-27 Alessandro Ragano , Emmanouil Benetos , Michael Chinen , Helard B. Martinez , Chandan K. A. Reddy , Jan Skoglund , Andrew Hines

While deep learning has made impressive progress in speech synthesis and voice conversion, the assessment of the synthesized speech is still carried out by human participants. Several recent papers have proposed deep-learning-based…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-10 Yeunju Choi , Youngmoon Jung , Hoirin Kim

This paper proposes an noise type classification aided attention-based neural network approach for monaural speech enhancement. The network is constructed based on a previous work by introducing a noise classification subnetwork into the…

Sound · Computer Science 2021-06-01 Lu Ma , Song Yang , Yaguang Gong , Zhongqin Wu

Although deep learning (DL) has achieved notable progress in speech enhancement (SE), further research is still required for a DL-based SE system to adapt effectively and efficiently to particular speakers. In this study, we propose a novel…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-11 Cheng Yu , Szu-Wei Fu , Tsun-An Hsieh , Yu Tsao , Mirco Ravanelli

Wideband codecs such as AMR-WB or EVS are widely used in (mobile) speech communication. Evaluation of coded speech quality is often performed subjectively by an absolute category rating (ACR) listening test. However, the ACR test is…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-20 Ziyi Xu , Ziyue Zhao , Tim Fingscheidt

Subjective listening tests remain the golden standard for speech quality assessment, but are costly, variable, and difficult to scale. In contrast, existing objective metrics, such as PESQ, F0 correlation, and DNSMOS, typically capture only…

Sound · Computer Science 2025-05-28 Jiatong Shi , Hye-Jin Shim , Shinji Watanabe

An effective approach to automatically predict the subjective rating for synthetic speech is to train on a listening test dataset with human-annotated scores. Although each speech sample in the dataset is rated by several listeners, most…

Sound · Computer Science 2021-10-19 Wen-Chin Huang , Erica Cooper , Junichi Yamagishi , Tomoki Toda

Speech quality assessment is a critical process in selecting text-to-speech synthesis (TTS) or voice conversion models. Evaluation of voice synthesis can be done using objective metrics or subjective metrics. Although there are many…

Sound · Computer Science 2025-06-04 Saurabh Agrawal , Raj Gohil , Gopal Kumar Agrawal , Vikram C M , Kushal Verma

Current state-of-the-art methods for automatic synthetic speech evaluation are based on MOS prediction neural models. Such MOS prediction models include MOSNet and LDNet that use spectral features as input, and SSL-MOS that relies on a…

Speech quality assessment (SQA) is often used to learn a mapping from a high-dimensional input space to a scalar that represents the mean opinion score (MOS) of the perceptual speech quality. Learning such a mapping is challenging for many…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-17 Junyi Fan , Donald Williamson

Objective speech quality models aim to predict human-perceived speech quality using automated methods. However, cross-lingual generalization remains a major challenge, as Mean Opinion Scores (MOS) vary across languages due to linguistic,…

Computation and Language · Computer Science 2025-02-19 Wafaa Wardah , Tuğçe Melike Koçak Büyüktaş , Kirill Shchegelskiy , Sebastian Möller , Robert P. Spang