English
Related papers

Related papers: NORESQA: A Framework for Speech Quality Assessment…

200 papers

Interacting with a speech interface to query a Question Answering (QA) system is becoming increasingly popular. Typically, QA systems rely on passage retrieval to select candidate contexts and reading comprehension to extract the final…

Computation and Language · Computer Science 2022-09-28 Georgios Sidiropoulos , Svitlana Vakulenko , Evangelos Kanoulas

Speech quality assessment has been a critical component in many voice communication related applications such as telephony and online conferencing. Traditional intrusive speech quality assessment requires the clean reference of the degraded…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-07 Yuchen Liu , Li-Chia Yang , Alex Pawlicki , Marko Stamenovic

Non-intrusive speech quality assessment (SQA) systems suffer from limited training data and costly human annotations, hindering their generalization to real-time conferencing calls. In this work, we propose leveraging large language models…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-11 Fredrik Cumlin , Xinyu Liang , Anubhab Ghosh , Saikat Chatterjee

Without the need for a clean reference, non-intrusive speech assessment methods have caught great attention for objective evaluations. While deep learning models have been used to develop non-intrusive speech assessment methods with…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-16 Hsin-Tien Chiang , Szu-Wei Fu , Hsin-Min Wang , Yu Tsao , John H. L. Hansen

Although recent neural text-to-speech (TTS) systems have achieved high-quality speech synthesis, there are cases where a TTS system generates low-quality speech, mainly caused by limited training data or information loss during knowledge…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-26 Yeunju Choi , Youngmoon Jung , Youngjoo Suh , Hoirin Kim

The quantification of audio aesthetics remains a complex challenge in audio processing, primarily due to its subjective nature, which is influenced by human perception and cultural context. Traditional methods often depend on human…

The evolution of point cloud processing algorithms necessitates an accurate assessment for their quality. Previous works consistently regard point cloud quality assessment (PCQA) as a MOS regression problem and devise a deterministic…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Songlin Fan , Wei Gao , Zhineng Chen , Ge Li , Guoqing Liu , Qicheng Wang

Background noise is a major source of quality impairments in Voice over Internet Protocol (VoIP) and Public Switched Telephone Network (PSTN) calls. Recent work shows the efficacy of deep learning for noise suppression, but the datasets…

An effective approach to automatically predict the subjective rating for synthetic speech is to train on a listening test dataset with human-annotated scores. Although each speech sample in the dataset is rated by several listeners, most…

Sound · Computer Science 2021-10-19 Wen-Chin Huang , Erica Cooper , Junichi Yamagishi , Tomoki Toda

Traditional image quality assessment (IQA) methods rely on mean opinion scores (MOS), which are resource-intensive to collect and fail to provide interpretable, localized feedback on specific image distortions. We overcome these limitations…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Fadeel Sher Khan , Long N. Le , Abhinau K. Venkataramanan , Seok-Jun Lee , Hamid R. Sheikh

The evaluation of synthetic and processed speech has long been a cornerstone of audio engineering and speech science. Although subjective listening tests remain the gold standard for assessing perceptual quality and intelligibility, their…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-05 Yu Tsao

The ground truth used for training image, video, or speech quality prediction models is based on the Mean Opinion Scores (MOS) obtained from subjective experiments. Usually, it is necessary to conduct multiple experiments, mostly with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-15 Gabriel Mittag , Saman Zadtootaghaj , Thilo Michael , Babak Naderi , Sebastian Möller

Source separation is a crucial pre-processing step for various speech processing tasks, such as automatic speech recognition (ASR). Traditionally, the evaluation metrics for speech separation rely on the matched reference audios and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-28 Ari Frummer , Helin Wang , Tianyu Cao , Adi Arbel , Yuval Sieradzki , Oren Gal , Jesús Villalba , Thomas Thebaud , Najim Dehak

Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjective assessments, the gold standard for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-28 Pablo M. Delgado , Jürgen Herre

Recruitment of appropriate people for certain positions is critical for any companies or organizations. Manually screening to select appropriate candidates from large amounts of resumes can be exhausted and time-consuming. However, there is…

Information Retrieval · Computer Science 2018-10-09 Yong Luo , Huaizheng Zhang , Yongjie Wang , Yonggang We , Xinwen Zhang

Perceptual speech quality is an important performance metric for teleconferencing applications. The mean opinion score (MOS) is standardized for the perceptual evaluation of speech quality and is obtained by asking listeners to rate the…

Sound · Computer Science 2022-12-06 Haleh Akrami , Hannes Gamper

Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep…

Computation and Language · Computer Science 2016-11-29 Brian Patton , Yannis Agiomyrgiannakis , Michael Terry , Kevin Wilson , Rif A. Saurous , D. Sculley

Research in modeling subjective metrics for quality assessment has led to the development of no-reference speech models that directly operate on utterance waveforms to predict the Mean Opinion Score (MOS). These models often rely on…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-29 Imran E Kibria , Donald S. Williamson

Generative models for image restoration, enhancement, and generation have significantly improved the quality of the generated images. Surprisingly, these models produce more pleasant images to the human eye than other methods, yet, they may…

Image and Video Processing · Electrical Eng. & Systems 2022-04-28 Marcos V. Conde , Maxime Burchi , Radu Timofte

Measuring quality and intelligibility of a speech signal is usually a critical step in development of speech processing systems. To enable this, a variety of metrics to measure quality and intelligibility under different assumptions have…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-05 Anurag Kumar , Ke Tan , Zhaoheng Ni , Pranay Manocha , Xiaohui Zhang , Ethan Henderson , Buye Xu