中文
相关论文

相关论文: Rethinking Mean Opinion Scores in Speech Quality A…

200 篇论文

The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in the mean opinion…

音频与语音处理 · 电气工程与系统科学 2021-04-06 Meng Yu , Chunlei Zhang , Yong Xu , Shixiong Zhang , Dong Yu

In speech generation tasks, human subjective ratings, usually referred to as the opinion score, are considered the "gold standard" for speech quality evaluation, with the mean opinion score (MOS) serving as the primary evaluation metric.…

声音 · 计算机科学 2024-06-21 Yuxun Tang , Jiatong Shi , Yuning Wu , Qin Jin

Several studies have proposed deep-learning-based models to predict the mean opinion score (MOS) of synthesized speech, showing the possibility of replacing human raters. However, inter- and intra-rater variability in MOSs makes it hard to…

音频与语音处理 · 电气工程与系统科学 2020-12-03 Yeunju Choi , Youngmoon Jung , Hoirin Kim

Speech quality assessment (SQA) refers to the evaluation of speech quality, and developing an accurate automatic SQA method that reflects human perception has become increasingly important, in order to keep up with the generative AI boom.…

声音 · 计算机科学 2025-08-29 Wen-Chin Huang

This paper proposes an approach to improve Non-Intrusive speech quality assessment(NI-SQA) based on the residuals between impaired speech and enhanced speech. The difficulty in our task is particularly lack of information, for which the…

声音 · 计算机科学 2022-03-23 Zhe Ye , Jiahao Chen , Diqun Yan

The rapid proliferation of AI-Generated Content (AIGC) has necessitated robust metrics for perceptual quality assessment. However, automatic Mean Opinion Score (MOS) prediction models are often compromised by data scarcity, predisposing…

音频与语音处理 · 电气工程与系统科学 2026-03-18 Kuan-Tang Huang , Chien-Chun Wang , Cheng-Yeh Yang , Hung-Shin Lee , Hsin-Min Wang , Berlin Chen

To compare the performance of two speech generation systems, one of the most effective approaches is estimating the preference score between their generated speech. This paper proposes a novel universal preference-score-based pairwise…

声音 · 计算机科学 2025-06-03 Yu-Fei Shi , Yang Ai , Zhen-Hua Ling

Assessing the naturalness of speech using mean opinion score (MOS) prediction models has positive implications for the automatic evaluation of speech synthesis systems. Early MOS prediction models took the raw waveform or amplitude spectrum…

声音 · 计算机科学 2024-11-19 Yu-Fei Shi , Yang Ai , Ye-Xin Lu , Hui-Peng Du , Zhen-Hua Ling

Modern speech quality prediction models are trained on audio data resampled to a specific sampling rate. When faced with higher-rate audio at test time, these models can produce biased scores. We introduce HighRateMOS, the first…

音频与语音处理 · 电气工程与系统科学 2025-06-30 Wenze Ren , Yi-Cheng Lin , Wen-Chin Huang , Ryandhimas E. Zezario , Szu-Wei Fu , Sung-Feng Huang , Erica Cooper , Haibin Wu , Hung-Yu Wei , Hsin-Min Wang , Hung-yi Lee , Yu Tsao

Accurate audio quality estimation is essential for developing and evaluating audio generation, retrieval, and enhancement systems. Existing non-intrusive assessment models predict a single Mean Opinion Score (MOS) for speech, merging…

音频与语音处理 · 电气工程与系统科学 2026-01-13 Yi-Cheng Lin , Jia-Hung Chen , Hung-yi Lee

Text-to-Speech synthesis systems are generally evaluated using Mean Opinion Score (MOS) tests, where listeners score samples of synthetic speech on a Likert scale. A major drawback of MOS tests is that they only offer a general measure of…

音频与语音处理 · 电气工程与系统科学 2021-07-07 Elijah Gutierrez , Pilar Oplustil-Gallegos , Catherine Lai

Non-intrusive speech quality assessment is a crucial operation in multimedia applications. The scarcity of annotated data and the lack of a reference signal represent some of the main challenges for designing efficient quality assessment…

音频与语音处理 · 电气工程与系统科学 2021-08-20 Alessandro Ragano , Emmanouil Benetos , Andrew Hines

Multimodal Sentiment Analysis (MSA) aims to infer human sentiment from textual, acoustic, and visual signals. In real-world scenarios, however, multimodal inputs are often compromised by dynamic noise or modality missingness. Existing…

人工智能 · 计算机科学 2026-04-09 Yitong Zhu , Yuxuan Jiang , Guanxuan Jiang , Bojing Hou , Peng Yuan Zhou , Ge Lin Kan , Yuyang Wang

In spoken question answering, the systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that human seek or test their knowledge is via human conversations.…

计算与语言 · 计算机科学 2022-05-02 Chenyu You , Nuo Chen , Fenglin Liu , Shen Ge , Xian Wu , Yuexian Zou

Assessing the perceptual quality of synthetic speech is crucial for guiding the development and refinement of speech generation models. However, it has traditionally relied on human subjective ratings such as the Mean Opinion Score (MOS),…

Recent speech enhancement (SE) models increasingly leverage self-supervised learning (SSL) representations for their rich semantic information. Typically, intermediate features are aggregated into a single representation via a lightweight…

声音 · 计算机科学 2026-02-02 Seungu Han , Sungho Lee , Kyogu Lee

In online conferencing applications, estimating the perceived quality of an audio signal is crucial to ensure high quality of experience for the end user. The most reliable way to assess the quality of a speech signal is through human…

音频与语音处理 · 电气工程与系统科学 2023-08-24 Bastiaan Tamm , Rik Vandenberghe , Hugo Van hamme

Existing objective evaluation metrics for voice conversion (VC) are not always correlated with human perception. Therefore, training VC models with such criteria may not effectively improve naturalness and similarity of converted speech. In…

声音 · 计算机科学 2022-03-01 Chen-Chou Lo , Szu-Wei Fu , Wen-Chin Huang , Xin Wang , Junichi Yamagishi , Yu Tsao , Hsin-Min Wang

Many computer scientists use the aggregated answers of online workers to represent ground truth. Prior work has shown that aggregation methods such as majority voting are effective for measuring relatively objective features. For subjective…

计算与语言 · 计算机科学 2021-04-06 Jiele Wu , Chau-Wai Wong , Xinyan Zhao , Xianpeng Liu

Previous methods for predicting room acoustic parameters and speech quality metrics have focused on the single-channel case, where room acoustics and Mean Opinion Score (MOS) are predicted for a single recording device. However,…

音频与语音处理 · 电气工程与系统科学 2024-03-14 Jozef Coldenhoff , Andrew Harper , Paul Kendrick , Tijana Stojkovic , Milos Cernak