English
Related papers

Related papers: Improving Self-Supervised Learning-based MOS Predi…

200 papers

Speech synthesis has come a long way as current text-to-speech (TTS) models can now generate natural human-sounding speech. However, most of the TTS research focuses on using adult speech data and there has been very limited work done on…

Sound · Computer Science 2022-04-05 Rishabh Jain , Mariam Yiwere , Dan Bigioi , Peter Corcoran , Horia Cucu

We present the UTokyo-SaruLab mean opinion score (MOS) prediction system submitted to VoiceMOS Challenge 2022. The challenge is to predict the MOS values of speech samples collected from previous Blizzard Challenges and Voice Conversion…

Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-training method that tailors the foundational Audio Large…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-12 Elizaveta Kostenok , Mathieu Salzmann , Milos Cernak

In this study, we propose a cross-domain multi-objective speech assessment model called MOSA-Net, which can estimate multiple speech assessment metrics simultaneously. Experimental results show that MOSA-Net can improve the linear…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-20 Ryandhimas E. Zezario , Szu-Wei Fu , Fei Chen , Chiou-Shann Fuh , Hsin-Min Wang , Yu Tsao

The perceptual task of speech quality assessment (SQA) is a challenging task for machines to do. Objective SQA methods that rely on the availability of the corresponding clean reference have been the primary go-to approaches for SQA.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-19 Pranay Manocha , Buye Xu , Anurag Kumar

We propose MambaRate, which predicts Mean Opinion Scores (MOS) with limited bias regarding the sampling rate of the waveform under evaluation. It is designed for Track 3 of the AudioMOS Challenge 2025, which focuses on predicting MOS for…

Recent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the deepest or final…

Sound · Computer Science 2023-06-27 George Close , William Ravenscroft , Thomas Hain , Stefan Goetze

In this research, we propose an architectural solution to implement the voice over IP (VoIP) service in campus environment network. Voice over IP (VoIP) technology has become a discussion issue for this time being. Today, the deployment of…

Networking and Internet Architecture · Computer Science 2009-06-05 Mohd Nazri Ismail

We train a MOS prediction model based on wav2vec 2.0 using the open-access data sets BVCC and SOMOS. Our test with neural TTS data in the low-resource language (LRL) West Frisian shows that pre-training on BVCC before fine-tuning on SOMOS…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-01 Phat Do , Matt Coler , Jelske Dijkstra , Esther Klabbers

Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objective metrics are frequently used, but rarely validated…

Sound · Computer Science 2026-03-03 Christoph Minixhofer , Ondrej Klejch , Peter Bell

The rank correlation coefficients and the ranked-based statistical tests (as a subset of non-parametric techniques) might be misleading when they are applied to subjectively collected opinion scores. Those techniques assume that the data is…

Multimedia · Computer Science 2020-10-01 Babak Naderi , Sebastian Möller

Human subjective evaluation is the gold standard to evaluate speech quality optimized for human perception. Perceptual objective metrics serve as a proxy for subjective scores. The conventional and widely used metrics require a reference…

Sound · Computer Science 2021-02-12 Chandan K A Reddy , Vishak Gopal , Ross Cutler

Objective estimators of multimedia quality are often judged by comparing estimates with subjective "truth data," most often via Pearson correlation coefficient (PCC) or mean-squared error (MSE). But subjective test results contain noise, so…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-16 Jaden Pieper , Stephen D. Voran

A common problem for automatic speech recognition systems is how to recognize words that they did not see during training. Currently there is no established method of evaluating different techniques for tackling this problem. We propose…

Computation and Language · Computer Science 2021-07-20 Rudolf A. Braun , Srikanth Madikeri , Petr Motlicek

The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in the mean opinion…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-06 Meng Yu , Chunlei Zhang , Yong Xu , Shixiong Zhang , Dong Yu

There has been a growing demand for automated spoken language assessment systems in recent years. A standard pipeline for this process is to start with a speech recognition system and derive features, either hand-crafted or based on…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-17 Stefano Bannò , Kate M. Knill , Marco Matassoni , Vyas Raina , Mark J. F. Gales

Methods for automatically assessing speech quality in real world environments are critical for developing robust human language technologies and assistive devices. Behavioral ratings provided by human raters (e.g., mean opinion scores; MOS)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-09 Mattson Ogg , Caitlyn Bishop , Han Yi , Sarah Robinson

Recent advancements in Deep and Self-Supervised Learning (SSL) have led to substantial improvements in Speech Emotion Recognition (SER) performance, reaching unprecedented levels. However, obtaining sufficient amounts of accurately labeled…

Computation and Language · Computer Science 2025-02-25 Bulat Khaertdinov , Pedro Jeuris , Annanda Sousa , Enrique Hortal

The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS…

Diffusion-based text-to-speech (TTS) systems have made remarkable progress in zero-shot speech synthesis, yet optimizing all components for perceptual metrics remains challenging. Prior work with DMOSpeech demonstrated direct metric…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-22 Yinghao Aaron Li , Xilin Jiang , Fei Tao , Cheng Niu , Kaifeng Xu , Juntong Song , Nima Mesgarani
‹ Prev 1 3 4 5 6 7 10 Next ›