中文
相关论文

相关论文: WhiSQA: Non-Intrusive Speech Quality Prediction Us…

200 篇论文

Vocoders, encoding speech signals into acoustic features and allowing for speech signal reconstruction from them, have been studied for decades. Recently, the rise of deep learning has particularly driven the development of neural vocoders…

音频与语音处理 · 电气工程与系统科学 2025-07-03 Shaowen Chen , Tomoki Toda

As speech processing systems in mobile and edge devices become more commonplace, the demand for unintrusive speech quality monitoring increases. Deep learning methods provide high-quality estimates of objective and subjective speech quality…

Automatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have achieved progress in…

声音 · 计算机科学 2025-04-30 Zhicheng Lian , Lizhi Wang , Hua Huang

We investigate the performance of features that can capture nonlinear recurrence dynamics embedded in the speech signal for the task of Speech Emotion Recognition (SER). Reconstruction of the phase space of each speech frame and the…

Speech severity evaluation is becoming increasingly important as the economic burden of speech disorders grows. Current speech severity models often struggle with generalization, learning dataset-specific acoustic cues rather than…

声音 · 计算机科学 2025-10-02 Bence Mark Halpern , Tomoki Toda

The evaluation of synthetic and processed speech has long been a cornerstone of audio engineering and speech science. Although subjective listening tests remain the gold standard for assessing perceptual quality and intelligibility, their…

音频与语音处理 · 电气工程与系统科学 2025-09-05 Yu Tsao

Transformer-based models have demonstrated their effectiveness in automatic speech recognition (ASR) tasks and even shown superior performance over the conventional hybrid framework. The main idea of Transformers is to capture the…

声音 · 计算机科学 2022-07-05 Kun Wei , Pengcheng Guo , Ning Jiang

Silent Speech Interfaces (SSIs) offer a noninvasive alternative to brain-computer interfaces for soundless verbal communication. We introduce Multimodal Orofacial Neural Audio (MONA), a system that leverages cross-modal alignment through…

人机交互 · 计算机科学 2024-03-12 Tyler Benster , Guy Wilson , Reshef Elisha , Francis R Willett , Shaul Druckmann

Text-based machine comprehension (MC) systems have a wide-range of applications, and standard corpora exist for developing and evaluating approaches. There has been far less research on spoken question answering (SQA) systems. The SQA task…

计算与语言 · 计算机科学 2021-07-13 Vatsal Raina , Mark J. F. Gales

The proposed model aims to develop a speech recognition technology for hearing, speech, or cognitively disabled people. All the available technology in the field of speech recognition doesn't come with an interface for communication for…

声音 · 计算机科学 2024-07-23 Ritabrata Roy Choudhury

Automatic speech recognition (ASR) has gained remarkable successes thanks to recent advances of deep learning, but it usually degrades significantly under real-world noisy conditions. Recent works introduce speech enhancement (SE) as…

音频与语音处理 · 电气工程与系统科学 2024-04-19 Yuchen Hu , Chen Chen , Qiushi Zhu , Eng Siong Chng

Neural network architectures are at the core of powerful automatic speech recognition systems (ASR). However, while recent researches focus on novel model architectures, the acoustic input features remain almost unchanged. Traditional ASR…

音频与语音处理 · 电气工程与系统科学 2018-11-27 Titouan Parcollet , Mirco Ravanelli , Mohamed Morchid , Georges Linarès , Renato De Mori

This paper presents a speech intelligibility model based on automatic speech recognition (ASR), combining phoneme probabilities from deep neural networks (DNN) and a performance measure that estimates the word error rate from these…

Large speech recognition models like Whisper-small achieve high accuracy but are difficult to deploy on edge devices due to their high computational demand. To this end, we present a unified, cross-library evaluation of post-training…

音频与语音处理 · 电气工程与系统科学 2026-05-22 Arthur Söhler , Julian Irigoyen , Andreas Søeborg Kirkedal

Automatic Speech Recognition (ASR) systems, despite large multilingual training, struggle in low-resource scenarios where labeled data is scarce. We propose BEARD (BEST-RQ Encoder Adaptation with Re-training and Distillation), a novel…

计算与语言 · 计算机科学 2026-05-13 Raphaël Bagat , Irina Illina , Emmanuel Vincent

Audio quality assessment is critical for assessing the perceptual realism of sounds. However, the time and expense of obtaining ''gold standard'' human judgments limit the availability of such data. For AR&VR, good perceived sound quality…

音频与语音处理 · 电气工程与系统科学 2022-06-27 Pranay Manocha , Anurag Kumar , Buye Xu , Anjali Menon , Israel D. Gebru , Vamsi K. Ithapu , Paul Calamia

Recent research on speech enhancement (SE) has seen the emergence of deep-learning-based methods. It is still a challenging task to determine the effective ways to increase the generalizability of SE under diverse test conditions. In this…

音频与语音处理 · 电气工程与系统科学 2021-09-01 Ryandhimas E. Zezario , Chiou-Shann Fuh , Hsin-Min Wang , Yu Tsao

This study addresses the speech enhancement (SE) task within the causal inference paradigm by modeling the noise presence as an intervention. Based on the potential outcome framework, the proposed causal inference-based speech enhancement…

音频与语音处理 · 电气工程与系统科学 2022-11-03 Tsun-An Hsieh , Chao-Han Huck Yang , Pin-Yu Chen , Sabato Marco Siniscalchi , Yu Tsao

Self-supervised learning (SSL) has grown in interest within the speech processing community, since it produces representations that are useful for many downstream tasks. SSL uses global and contextual methods to produce robust…

音频与语音处理 · 电气工程与系统科学 2024-11-08 Subrina Sultana , Donald S. Williamson

Deep neural networks have largely demonstrated their ability to perform automated speech recognition (ASR) by extracting meaningful features from input audio frames. Such features, however, may consist not only of information about the…

音频与语音处理 · 电气工程与系统科学 2022-09-16 David M. Chan , Shalini Ghosh
‹ 上一页 1 8 9 10 下一页 ›