中文
相关论文

相关论文: Improving Self-Supervised Learning-based MOS Predi…

200 篇论文

We aim to characterize how different speakers contribute to the perceived output quality of multi-speaker Text-to-Speech (TTS) synthesis. We automatically rate the quality of TTS using a neural network (NN) trained on human mean opinion…

计算与语言 · 计算机科学 2020-04-28 Jennifer Williams , Joanna Rownicka , Pilar Oplustil , Simon King

Modern speech quality prediction models are trained on audio data resampled to a specific sampling rate. When faced with higher-rate audio at test time, these models can produce biased scores. We introduce HighRateMOS, the first…

音频与语音处理 · 电气工程与系统科学 2025-06-30 Wenze Ren , Yi-Cheng Lin , Wen-Chin Huang , Ryandhimas E. Zezario , Szu-Wei Fu , Sung-Feng Huang , Erica Cooper , Haibin Wu , Hung-Yu Wei , Hsin-Min Wang , Hung-yi Lee , Yu Tsao

In this work, we present the SOMOS dataset, the first large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples. It can be employed to train automatic MOS prediction systems focused on the…

Speech quality in online conferencing applications is typically assessed through human judgements in the form of the mean opinion score (MOS) metric. Since such a labor-intensive approach is not feasible for large-scale speech quality…

音频与语音处理 · 电气工程与系统科学 2022-10-04 Bastiaan Tamm , Helena Balabin , Rik Vandenberghe , Hugo Van hamme

Prosody is essential for speech technology, shaping comprehension, naturalness, and expressiveness. However, current text-to-speech (TTS) systems still struggle to accurately capture human-like prosodic variation, in part because existing…

音频与语音处理 · 电气工程与系统科学 2025-11-05 Cedric Chan , Jianjing Kuang

Speech quality assessment (SQA) aims to evaluate the quality of speech samples without relying on time-consuming listener questionnaires. Recent efforts have focused on training neural-based SQA models to predict the mean opinion score…

声音 · 计算机科学 2025-06-24 Yuto Kondo , Hirokazu Kameoka , Kou Tanaka , Takuhiro Kaneko

This paper investigates the use of Mean Opinion Score (MOS), a common image quality metric, as a user-centric evaluation metric for XAI post-hoc explainers. To measure the MOS, a user experiment is proposed, which has been conducted with…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Hyeon Yu , Jenny Benois-Pineau , Romain Bourqui , Romain Giot , Alexey Zhukov

Speech evaluation measures a learners oral proficiency using automatic models. Corpora for training such models often pose sparsity challenges given that there often is limited scored data from teachers, in addition to the score…

人工智能 · 计算机科学 2024-09-24 Huayun Zhang , Jeremy H. M. Wong , Geyu Lin , Nancy F. Chen

Automatic speech quality assessment has become increasingly important as modern speech generation systems continue to advance, while human listening tests remain costly, time-consuming, and difficult to scale. Most existing learning-based…

Speech quality assessment has been a critical component in many voice communication related applications such as telephony and online conferencing. Traditional intrusive speech quality assessment requires the clean reference of the degraded…

音频与语音处理 · 电气工程与系统科学 2022-11-07 Yuchen Liu , Li-Chia Yang , Alex Pawlicki , Marko Stamenovic

Automatic Mean Opinion Score (MOS) prediction is crucial to evaluate the perceptual quality of the synthetic speech. While recent approaches using pre-trained self-supervised learning (SSL) models have shown promising results, they only…

音频与语音处理 · 电气工程与系统科学 2023-09-01 Hui Wang , Shiwan Zhao , Xiguang Zheng , Yong Qin

Automatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have achieved progress in…

声音 · 计算机科学 2025-04-30 Zhicheng Lian , Lizhi Wang , Hua Huang

We present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tracks. The first track was on predicting the quality of…

This paper introduces a novel objective function for quality mean opinion score (MOS) prediction of unseen speech synthesis systems. The proposed function measures the similarity of relative positions of predicted MOS values, in a…

音频与语音处理 · 电气工程与系统科学 2023-10-10 Hemant Yadav , Erica Cooper , Junichi Yamagishi , Sunayana Sitaram , Rajiv Ratn Shah

In this paper, we propose a highly efficient method to estimate an image's mean opinion score (MOS) from a single opinion score (SOS). Assuming that each SOS is the observed sample of a normal distribution and the MOS is its unknown…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Lei Wang , Desen Yuan

One objective of Speech Quality Assessment (SQA) is to estimate the ranks of synthetic speech systems. However, recent SQA models are typically trained using low-precision direct scores such as mean opinion scores (MOS) as the training…

音频与语音处理 · 电气工程与系统科学 2023-08-30 Cheng-Hung Hu , Yusuke Yasuda , Tomoki Toda

With the advancement of self-supervised learning (SSL), fine-tuning pretrained SSL models for mean opinion score (MOS) prediction has achieved state-of-the-art performance. However, during fine-tuning, these SSL-based MOS prediction models…

声音 · 计算机科学 2026-01-21 Jianing Yang , Wataru Nakata , Yuki Saito , Hiroshi Saruwatari

Previous methods for predicting room acoustic parameters and speech quality metrics have focused on the single-channel case, where room acoustics and Mean Opinion Score (MOS) are predicted for a single recording device. However,…

音频与语音处理 · 电气工程与系统科学 2024-03-14 Jozef Coldenhoff , Andrew Harper , Paul Kendrick , Tijana Stojkovic , Milos Cernak

The field of prosody transfer in speech synthesis systems is rapidly advancing. This research is focused on evaluating learning methods for adapting pre-trained monolingual text-to-speech (TTS) models to multilingual conditions, i.e.,…

计算与语言 · 计算机科学 2024-06-19 Arnav Goel , Medha Hira , Anubha Gupta

We present the second edition of the VoiceMOS Challenge, a scientific event that aims to promote the study of automatic prediction of the mean opinion score (MOS) of synthesized and processed speech. This year, we emphasize real-world and…

音频与语音处理 · 电气工程与系统科学 2023-10-10 Erica Cooper , Wen-Chin Huang , Yu Tsao , Hsin-Min Wang , Tomoki Toda , Junichi Yamagishi