中文
相关论文

相关论文: Coded Speech Quality Measurement by a Non-Intrusiv…

200 篇论文

The common standard for quality evaluation of automatic speech recognition (ASR) systems is reference-based metrics such as the Word Error Rate (WER), computed using manual ground-truth transcriptions that are time-consuming and expensive…

计算与语言 · 计算机科学 2023-06-26 Kamer Ali Yuksel , Thiago Ferreira , Ahmet Gunduz , Mohamed Al-Badrashiny , Golara Javadi

In the realm of automatic speech recognition (ASR), the quest for models that not only perform with high accuracy but also offer transparency in their decision-making processes is crucial. The potential of quality estimation (QE) metrics is…

计算与语言 · 计算机科学 2024-02-06 Golara Javadi , Kamer Ali Yuksel , Yunsu Kim , Thiago Castro Ferreira , Mohamed Al-Badrashiny

This paper introduces HAAQI-Net, a non-intrusive deep learning-based music audio quality assessment model for hearing aid users. Unlike traditional methods like the Hearing Aid Audio Quality Index (HAAQI) that require intrusive reference…

音频与语音处理 · 电气工程与系统科学 2025-01-10 Dyah A. M. G. Wisnu , Stefano Rini , Ryandhimas E. Zezario , Hsin-Min Wang , Yu Tsao

Measuring performance of an automatic speech recognition (ASR) system without ground-truth could be beneficial in many scenarios, especially with data from unseen domains, where performance can be highly inconsistent. In conventional ASR…

计算与语言 · 计算机科学 2019-04-11 Ruizhi Li , Gregory Sell , Hynek Hermansky

Text encodings from automatic speech recognition (ASR) transcripts and audio representations have shown promise in speech emotion recognition (SER) ever since. Yet, it is challenging to explain the effect of each information stream on the…

Training of speech enhancement systems often does not incorporate knowledge of human perception and thus can lead to unnatural sounding results. Incorporating psychoacoustically motivated speech perception metrics as part of model training…

声音 · 计算机科学 2022-06-16 George Close , Thomas Hain , Stefan Goetze

Non-intrusive intelligibility prediction is important for its application in realistic scenarios, where a clean reference signal is difficult to access. The construction of many non-intrusive predictors require either ground truth…

音频与语音处理 · 电气工程与系统科学 2022-07-07 Zehai Tu , Ning Ma , Jon Barker

In this paper we propose to exploit the automatic Quality Estimation (QE) of ASR hypotheses to perform the unsupervised adaptation of a deep neural network modeling acoustic probabilities. Our hypothesis is that significant improvements can…

计算与语言 · 计算机科学 2017-02-07 Daniele Falavigna , Marco Matassoni , Shahab Jalalvand , Matteo Negri , Marco Turchi

Neural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance downstream training efficiency and compatibility with…

We propose a training method for deep neural network (DNN)-based source enhancement to increase objective sound quality assessment (OSQA) scores such as the perceptual evaluation of speech quality (PESQ). In many conventional studies, DNNs…

机器学习 · 统计学 2018-10-23 Yuma Koizumi , Kenta Niwa , Yusuke Hioka , Kazunori Kobayashi , Yoichi Haneda

Neural networks have proven to be a formidable tool to tackle the problem of speech coding at very low bit rates. However, the design of a neural coder that can be operated robustly under real-world conditions remains a major challenge.…

音频与语音处理 · 电气工程与系统科学 2022-07-08 Nicola Pia , Kishan Gupta , Srikanth Korse , Markus Multrus , Guillaume Fuchs

We participated in track 2 of the VoiceMOS Challenge 2024, which aimed to predict the mean opinion score (MOS) of singing samples. Our submission secured the first place among all participating teams, excluding the official baseline. In…

声音 · 计算机科学 2024-12-24 Yu-Fei Shi , Yang Ai , Ye-Xin Lu , Hui-Peng Du , Zhen-Hua Ling

Subjective listening tests remain the golden standard for speech quality assessment, but are costly, variable, and difficult to scale. In contrast, existing objective metrics, such as PESQ, F0 correlation, and DNSMOS, typically capture only…

声音 · 计算机科学 2025-05-28 Jiatong Shi , Hye-Jin Shim , Shinji Watanabe

The goal in speech enhancement is to obtain an estimate of clean speech starting from the noisy signal by minimizing a chosen distortion measure, which results in an estimate that depends on the unknown clean signal or its statistics. Since…

音频与语音处理 · 电气工程与系统科学 2017-10-12 Jishnu Sadasivan , Chandra Sekhar Seelamantula , Nagarjuna Reddy Muraka

The estimation of speech intelligibility is still far from being a solved problem. Especially one aspect is problematic: most of the standard models require a clean reference signal in order to estimate intelligibility. This is an issue of…

音频与语音处理 · 电气工程与系统科学 2021-10-29 Mahdie Karbasi , Stefan Bleeck , Dorothea Kolossa

Automatic subjective speech quality assessment (SSQA) traditionally estimates speech quality on an utterance or system level. While this resolution was adequate for older transmission or synthesis systems that produced speech signals of…

音频与语音处理 · 电气工程与系统科学 2026-05-21 Michael Kuhlmann , Tobias Cord-Landwehr , Reinhold Haeb-Umbach

Automatic speech recognition (ASR) has gained remarkable successes thanks to recent advances of deep learning, but it usually degrades significantly under real-world noisy conditions. Recent works introduce speech enhancement (SE) as…

音频与语音处理 · 电气工程与系统科学 2024-04-19 Yuchen Hu , Chen Chen , Qiushi Zhu , Eng Siong Chng

The potential of deep learning in clinical speech processing is immense, yet the hurdles of limited and imbalanced clinical data samples loom large. This article addresses these challenges by showcasing the utilization of automatic speech…

Over the past few decades, computational methods have been developed to estimate perceptual audio quality. These methods, also referred to as objective quality measures, are usually developed and intended for a specific application domain.…

音频与语音处理 · 电气工程与系统科学 2021-10-25 Matteo Torcoli , Thorsten Kastner , Jürgen Herre

Recent advances in speech foundation models are largely driven by scaling both model size and data, enabling them to perform a wide range of tasks, including speech recognition. Traditionally, ASR models are evaluated using metrics like…

计算与语言 · 计算机科学 2025-06-06 Abdul Waheed , Hanin Atwany , Rita Singh , Bhiksha Raj