English
Related papers

Related papers: CCATMos: Convolutional Context-aware Transformer N…

200 papers

In this study, we propose a cross-domain multi-objective speech assessment model called MOSA-Net, which can estimate multiple speech assessment metrics simultaneously. Experimental results show that MOSA-Net can improve the linear…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-20 Ryandhimas E. Zezario , Szu-Wei Fu , Fei Chen , Chiou-Shann Fuh , Hsin-Min Wang , Yu Tsao

As a subjective metric to evaluate the quality of synthesized speech, Mean opinion score~(MOS) usually requires multiple annotators to score the same speech. Such an annotation approach requires a lot of manpower and is also time-consuming.…

Sound · Computer Science 2023-06-21 Kexin Wang , Yunlong Zhao , Qianqian Dong , Tom Ko , Mingxuan Wang

The ground truth used for training image, video, or speech quality prediction models is based on the Mean Opinion Scores (MOS) obtained from subjective experiments. Usually, it is necessary to conduct multiple experiments, mostly with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-15 Gabriel Mittag , Saman Zadtootaghaj , Thilo Michael , Babak Naderi , Sebastian Möller

Speech classification tasks often require powerful language understanding models to grasp useful features, which becomes problematic when limited training data is available. To attain superior classification performance, we propose to…

Computation and Language · Computer Science 2024-07-26 Nicolae-Catalin Ristea , Andrei Anghel , Radu Tudor Ionescu

This position paper argues that Mean Opinion Score (MOS), while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. MOS reduces rich, context-sensitive human judgments…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Mohamed Amine Kerkouri , Marouane Tliba , Aladine Chetouani , Nour Aburaed , Alessandro Bruno

Predicting audio quality in voice synthesis and conversion systems is a critical yet challenging task, especially when traditional methods like Mean Opinion Scores (MOS) are cumbersome to collect at scale. This paper addresses the gap in…

Sound · Computer Science 2023-12-27 Aditya Ravuri , Erica Cooper , Junichi Yamagishi

Training of multi-speaker text-to-speech (TTS) systems relies on curated datasets based on high-quality recordings or audiobooks. Such datasets often lack speaker diversity and are expensive to collect. As an alternative, recent studies…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-13 Sewade Ogun , Vincent Colotte , Emmanuel Vincent

Text-to-Speech synthesis systems are generally evaluated using Mean Opinion Score (MOS) tests, where listeners score samples of synthetic speech on a Likert scale. A major drawback of MOS tests is that they only offer a general measure of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-07 Elijah Gutierrez , Pilar Oplustil-Gallegos , Catherine Lai

In online conferencing applications, estimating the perceived quality of an audio signal is crucial to ensure high quality of experience for the end user. The most reliable way to assess the quality of a speech signal is through human…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-24 Bastiaan Tamm , Rik Vandenberghe , Hugo Van hamme

End-to-end (E2E) automatic speech recognition (ASR) systems often have difficulty recognizing uncommon words, that appear infrequently in the training data. One promising method, to improve the recognition accuracy on such rare words, is to…

Computation and Language · Computer Science 2021-11-08 Feng-Ju Chang , Jing Liu , Martin Radfar , Athanasios Mouchtaris , Maurizio Omologo , Ariya Rastrow , Siegfried Kunzmann

Ensuring that Text-to-Speech (TTS) systems deliver human-perceived quality at scale is a central challenge for modern speech technologies. Human subjective evaluation protocols such as Mean Opinion Score (MOS) and Side-by-Side (SBS)…

Computation and Language · Computer Science 2026-04-13 Ilya Trofimenko , David Kocharyan , Aleksandr Zaitsev , Pavel Repnikov , Mark Levin , Nikita Shevtsov

We present a system for non-intrusive prediction of speech quality in noisy and enhanced speech, developed for Track 3 of the VoiceMOS 2024 Challenge. The task required estimating the ITU-T P.835 metrics SIG, BAK, and OVRL without reference…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-28 Marie Kunešová , Aleš Pražák , Jan Lehečka

Automatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have achieved progress in…

Sound · Computer Science 2025-04-30 Zhicheng Lian , Lizhi Wang , Hua Huang

We present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tracks. The first track was on predicting the quality of…

Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD), as we expect that MOS can be used to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-26 Wangjin Zhou , Zhengdong Yang , Chenhui Chu , Sheng Li , Raj Dabre , Yi Zhao , Tatsuya Kawahara

The quantification of audio aesthetics remains a complex challenge in audio processing, primarily due to its subjective nature, which is influenced by human perception and cultural context. Traditional methods often depend on human…

Diffusion models have found great success in generating high quality, natural samples of speech, but their potential for density estimation for speech has so far remained largely unexplored. In this work, we leverage an unconditional…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-16 Danilo de Oliveira , Julius Richter , Jean-Marie Lemercier , Simon Welker , Timo Gerkmann

Objective speech quality assessment is central to telephony, VoIP, and streaming systems, where large volumes of degraded audio must be monitored and optimized at scale. Classical metrics such as PESQ and POLQA approximate human mean…

Sound · Computer Science 2025-12-10 Mahathir Monjur , Shahriar Nirjon

Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjective assessments, the gold standard for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-28 Pablo M. Delgado , Jürgen Herre

In this work, we present the SOMOS dataset, the first large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples. It can be employed to train automatic MOS prediction systems focused on the…