English
Related papers

Related papers: Predicting score distribution to improve non-intru…

200 papers

Estimating the perceived quality of an audio signal is critical for many multimedia and audio processing systems. Providers strive to offer optimal and reliable services in order to increase the user quality of experience (QoE). In this…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-19 Anderson R. Avila , Hannes Gamper , Chandan Reddy , Ross Cutler , Ivan Tashev , Johannes Gehrke

Background noise is a major source of quality impairments in Voice over Internet Protocol (VoIP) and Public Switched Telephone Network (PSTN) calls. Recent work shows the efficacy of deep learning for noise suppression, but the datasets…

MOS (Mean Opinion Score) is a subjective method used for the evaluation of a system's quality. Telecommunications (for voice and video), and speech synthesis systems (for generated speech) are a few of the many applications of the method.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-26 Bálint Gyires-Tóth , Csaba Zainkó

Speech enhancement employing deep neural networks (DNNs) for denoising are called deep noise suppression (DNS). During training, DNS methods are typically trained with mean squared error (MSE) type loss functions, which do not guarantee…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-09 Ziyi Xu , Maximilian Strake , Tim Fingscheidt

Perceptual speech quality is an important performance metric for teleconferencing applications. The mean opinion score (MOS) is standardized for the perceptual evaluation of speech quality and is obtained by asking listeners to rate the…

Sound · Computer Science 2022-12-06 Haleh Akrami , Hannes Gamper

Mean opinion score (MOS) is a typical subjective evaluation metric for speech synthesis systems. Since collecting MOS is time-consuming, it would be desirable if there are accurate MOS prediction models for automatic evaluation. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-16 Wei-Cheng Tseng , Wei-Tsung Kao , Hung-yi Lee

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020. We open…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-28 Chandan K A Reddy , Harishchandra Dubey , Vishak Gopal , Ross Cutler , Sebastian Braun , Hannes Gamper , Robert Aichner , Sriram Srinivasan

Perceptually-inspired objective functions such as the perceptual evaluation of speech quality (PESQ), signal-to-distortion ratio (SDR), and short-time objective intelligibility (STOI), have recently been used to optimize performance of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-27 Khandokar Md. Nayem , Donald S. Williamson

The INTERSPEECH 2020 Deep Noise Suppression (DNS) Challenge is intended to promote collaborative research in real-time single-channel Speech Enhancement aimed to maximize the subjective (perceptual) quality of the enhanced speech. A typical…

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP…

Human subjective evaluation is the gold standard to evaluate speech quality optimized for human perception. Perceptual objective metrics serve as a proxy for subjective scores. The conventional and widely used metrics require a reference…

Sound · Computer Science 2021-02-12 Chandan K A Reddy , Vishak Gopal , Ross Cutler

Human subjective evaluation is the gold standard to evaluate speech quality optimized for human perception. Perceptual objective metrics serve as a proxy for subjective scores. We have recently developed a non-intrusive speech quality…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-07 Chandan K A Reddy , Vishak Gopal , Ross Cutler

Mean opinion score (MOS) is a popular subjective metric to assess the quality of synthesized speech, and usually involves multiple human judges to evaluate each speech utterance. To reduce the labor cost in MOS test, multiple methods have…

Sound · Computer Science 2021-03-02 Yichong Leng , Xu Tan , Sheng Zhao , Frank Soong , Xiang-Yang Li , Tao Qin

Human judgments obtained through Mean Opinion Scores (MOS) are the most reliable way to assess the quality of speech signals. However, several recent attempts to automatically estimate MOS using deep learning approaches lack robustness and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-27 Pranay Manocha , Anurag Kumar

The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in the mean opinion…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-06 Meng Yu , Chunlei Zhang , Yong Xu , Shixiong Zhang , Dong Yu

An effective approach to automatically predict the subjective rating for synthetic speech is to train on a listening test dataset with human-annotated scores. Although each speech sample in the dataset is rated by several listeners, most…

Sound · Computer Science 2021-10-19 Wen-Chin Huang , Erica Cooper , Junichi Yamagishi , Tomoki Toda

Speech enhancement algorithms based on deep learning have been improved in terms of speech intelligibility and perceptual quality greatly. Many methods focus on enhancing the amplitude spectrum while reconstructing speech using the mixture…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-10 Qinglong Li , Fei Gao , Haixin Guan , Kaichi Ma

Existing objective evaluation metrics for voice conversion (VC) are not always correlated with human perception. Therefore, training VC models with such criteria may not effectively improve naturalness and similarity of converted speech. In…

Sound · Computer Science 2022-03-01 Chen-Chou Lo , Szu-Wei Fu , Wen-Chin Huang , Xin Wang , Junichi Yamagishi , Yu Tsao , Hsin-Min Wang

Modern speech quality prediction models are trained on audio data resampled to a specific sampling rate. When faced with higher-rate audio at test time, these models can produce biased scores. We introduce HighRateMOS, the first…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-30 Wenze Ren , Yi-Cheng Lin , Wen-Chin Huang , Ryandhimas E. Zezario , Szu-Wei Fu , Sung-Feng Huang , Erica Cooper , Haibin Wu , Hung-Yu Wei , Hsin-Min Wang , Hung-yi Lee , Yu Tsao

Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep…

Computation and Language · Computer Science 2016-11-29 Brian Patton , Yannis Agiomyrgiannakis , Michael Terry , Kevin Wilson , Rif A. Saurous , D. Sculley
‹ Prev 1 2 3 10 Next ›