中文
相关论文

相关论文: SCOREQ: Speech Quality Assessment with Contrastive…

200 篇论文

Speech quality assessment (SQA) aims to evaluate the quality of speech samples without relying on time-consuming listener questionnaires. Recent efforts have focused on training neural-based SQA models to predict the mean opinion score…

声音 · 计算机科学 2025-06-24 Yuto Kondo , Hirokazu Kameoka , Kou Tanaka , Takuhiro Kaneko

The recently proposed Sequence-to-Sequence (seq2seq) framework advocates replacing complex data processing pipelines, such as an entire automatic speech recognition system, with a single neural network trained in an end-to-end fashion. In…

神经与进化计算 · 计算机科学 2016-12-09 Jan Chorowski , Navdeep Jaitly

In this paper, we propose the coarse-to-fine optimization for the task of speech enhancement. Cosine similarity loss [1] has proven to be an effective metric to measure similarity of speech signals. However, due to the large variance of the…

声音 · 计算机科学 2019-08-23 Jian Yao , Ahmad Al-Dahle

Objective speech quality assessment is central to telephony, VoIP, and streaming systems, where large volumes of degraded audio must be monitored and optimized at scale. Classical metrics such as PESQ and POLQA approximate human mean…

声音 · 计算机科学 2025-12-10 Mahathir Monjur , Shahriar Nirjon

The diverse perceptual consequences of hearing loss severely impede speech communication, but standard clinical audiometry, which is focused on threshold-based frequency sensitivity, does not adequately capture deficits in frequency and…

音频与语音处理 · 电气工程与系统科学 2025-07-31 Xiajie Zhou , Candy Olivia Mawalim , Masashi Unoki

The category of figurative language contains many varieties, some of which are non-compositional in nature. This type of phrase or multi-word expression (MWE) includes idioms, which represent a single meaning that does not consist of the…

计算与语言 · 计算机科学 2026-03-25 Blake Matheny , Phuong Minh Nguyen , Minh Le Nguyen

Recent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the deepest or final…

声音 · 计算机科学 2023-06-27 George Close , William Ravenscroft , Thomas Hain , Stefan Goetze

Speech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and objective metrics. For instance, metrics like PESQ (Perceptual…

The mean opinion score (MOS) is a standard metric for assessing speech quality, but its singular focus fails to identify specific distortions when low scores are observed. The NISQA dataset addresses this limitation by providing ratings…

音频与语音处理 · 电气工程与系统科学 2025-06-06 Fredrik Cumlin , Xinyu Liang , Victor Ungureanu , Chandan K. A. Reddy , Christian Schüldt , Saikat Chatterjee

Speech evaluation measures a learners oral proficiency using automatic models. Corpora for training such models often pose sparsity challenges given that there often is limited scored data from teachers, in addition to the score…

人工智能 · 计算机科学 2024-09-24 Huayun Zhang , Jeremy H. M. Wong , Geyu Lin , Nancy F. Chen

Deep learning technology has been widely applied to speech enhancement. While testing the effectiveness of various network structures, researchers are also exploring the improvement of the loss function used in network training. Although…

音频与语音处理 · 电气工程与系统科学 2023-04-25 Tianrui Wang , Weibin Zhu

There has been significant research effort developing neural-network-based predictors of SQ in recent years. While a primary objective has been to develop non-intrusive, i.e.~reference-free, metrics to assess the performance of SE systems,…

声音 · 计算机科学 2025-08-05 George Close , Kris Hong , Thomas Hain , Stefan Goetze

Evaluating the quality of generated text automatically remains a significant challenge. Conventional reference-based metrics have been shown to exhibit relatively weak correlation with human evaluations. Recent research advocates the use of…

计算与语言 · 计算机科学 2025-11-25 Xiao Wang , Daniil Larionov , Siwei Wu , Yiqi Liu , Steffen Eger , Nafise Sadat Moosavi , Chenghua Lin

In this paper, we study the task of subjective speech quality assessment (SSQA), which refers to predicting the perceptual quality of speech. Owing to the development of deep neural network models, SSQA has greatly advanced and has been…

声音 · 计算机科学 2026-04-27 Wen-Chin Huang , Erica Cooper , Tomoki Toda

This paper introduces a novel objective function for quality mean opinion score (MOS) prediction of unseen speech synthesis systems. The proposed function measures the similarity of relative positions of predicted MOS values, in a…

音频与语音处理 · 电气工程与系统科学 2023-10-10 Hemant Yadav , Erica Cooper , Junichi Yamagishi , Sunayana Sitaram , Rajiv Ratn Shah

We propose a novel language-universal approach to end-to-end automatic spoken keyword recognition (SKR) leveraging upon (i) a self-supervised pre-trained model, and (ii) a set of universal speech attributes (manner and place of…

音频与语音处理 · 电气工程与系统科学 2024-06-05 Hao Yen , Pin-Jui Ku , Sabato Marco Siniscalchi , Chin-Hui Lee

We consider the problem of obtaining image quality representations in a self-supervised manner. We use prediction of distortion type and degree as an auxiliary task to learn features from an unlabeled image dataset containing a mixture of…

计算机视觉与模式识别 · 计算机科学 2022-06-30 Pavan C. Madhusudana , Neil Birkbeck , Yilin Wang , Balu Adsumilli , Alan C. Bovik

In this paper, we propose a global monotonicity consistency training strategy for quality assessment, which includes a differentiable, low-computation monotonicity evaluation loss function and a global perception training mechanism.…

图像与视频处理 · 电气工程与系统科学 2025-01-28 Yipeng Liu , Qi Yang , Yiling Xu

Speech emotion recognition is an important aspect of human-computer interaction. Prior work proposes various end-to-end models to improve the classification performance. However, most of them rely on the cross-entropy loss together with…

计算机视觉与模式识别 · 计算机科学 2019-10-25 Zheng Lian , Ya Li , Jianhua Tao , Jian Huang

The rapid proliferation of AI-Generated Content (AIGC) has necessitated robust metrics for perceptual quality assessment. However, automatic Mean Opinion Score (MOS) prediction models are often compromised by data scarcity, predisposing…

音频与语音处理 · 电气工程与系统科学 2026-03-18 Kuan-Tang Huang , Chien-Chun Wang , Cheng-Yeh Yang , Hung-Shin Lee , Hsin-Min Wang , Berlin Chen