中文
相关论文

相关论文: STOI-Net: A Deep Learning based Non-Intrusive Spee…

200 篇论文

Today's Automatic Speech Recognition systems only rely on acoustic signals and often don't perform well under noisy conditions. Performing multi-modal speech recognition - processing acoustic speech signals and lip-reading video…

计算机视觉与模式识别 · 计算机科学 2018-03-14 Matthijs Van keirsbilck , Bert Moons , Marian Verhelst

Controllable generation requires language models to realize output characteristics such as reading level, politeness, and toxicity. Existing steering methods are often indirect, require access to internal activations, or depend on auxiliary…

计算与语言 · 计算机科学 2026-05-29 Hyeseon An , Shinwoo Park , Hyundong Jin , Yo-Sub Han

Cochlear implants (CIs) are surgically implanted hearing devices, which allow to restore a sense of hearing in people suffering from profound hearing loss. Wireless streaming of audio from external devices to CI signal processors has become…

声音 · 计算机科学 2025-06-03 Reemt Hinrichs , Jörn Ostermann

Emotion classification of speech and assessment of the emotion strength are required in applications such as emotional text-to-speech and voice conversion. The emotion attribute ranking function based on Support Vector Machine (SVM) was…

声音 · 计算机科学 2022-06-16 Rui Liu , Berrak Sisman , Björn Schuller , Guanglai Gao , Haizhou Li

Although numerous recent studies have suggested new frameworks for zero-shot TTS using large-scale, real-world data, studies that focus on the intelligibility of zero-shot TTS are relatively scarce. Zero-shot TTS demands additional efforts…

音频与语音处理 · 电气工程与系统科学 2024-01-31 Sunghee Jung , Won Jang , Jaesam Yoon , Bongwan Kim

Speech data has rich acoustic and paralinguistic information with important cues for understanding a speaker's tone, emotion, and intent, yet traditional large language models such as BERT do not incorporate this information. There has been…

计算与语言 · 计算机科学 2023-11-14 Fatema Hasan , Yulong Li , James Foulds , Shimei Pan , Bishwaranjan Bhattacharjee

In this article, we provide a model to estimate a real-valued measure of the intelligibility of individual speech segments. We trained regression models based on Convolutional Neural Networks (CNN) for stop consonants…

音频与语音处理 · 电气工程与系统科学 2021-03-29 Ali Abavisani , Mark Hasegawa-Johnson

Supervised learning based methods for source localization, being data driven, can be adapted to different acoustic conditions via training and have been shown to be robust to adverse acoustic environments. In this paper, a convolutional…

音频与语音处理 · 电气工程与系统科学 2019-05-22 Soumitro Chakrabarty , Emanuël A. P. Habets

STOI-optimal masking has been previously proposed and developed for single-channel speech enhancement. In this paper, we consider the extension to the task of binaural speech enhancement in which spatial information is known to be important…

音频与语音处理 · 电气工程与系统科学 2022-10-03 Vikas Tokala , Mike Brookes , Patrick A. Naylor

Spoken language understanding (SLU) system usually consists of various pipeline components, where each component heavily relies on the results of its upstream ones. For example, Intent detection (ID), and slot filling (SF) require its…

计算与语言 · 计算机科学 2021-04-14 Di Wu , Yiren Chen , Liang Ding , Dacheng Tao

This paper considers speech enhancement of signals picked up in one noisy environment which must be presented to a listener in another noisy environment. Recently, it has been shown that an optimal solution to this problem requires the…

音频与语音处理 · 电气工程与系统科学 2022-05-06 Andreas Jonas Fuglsig , Jan Østergaard , Jesper Jensen , Lars Søndergaard Bertelsen , Peter Mariager , Zheng-Hua Tan

Recently, a semi-supervised learning method known as "noisy student training" has been shown to improve image classification performance of deep networks significantly. Noisy student training is an iterative self-training method that…

音频与语音处理 · 电气工程与系统科学 2020-11-02 Daniel S. Park , Yu Zhang , Ye Jia , Wei Han , Chung-Cheng Chiu , Bo Li , Yonghui Wu , Quoc V. Le

Intent classification is a fundamental task in the spoken language understanding field that has recently gained the attention of the scientific community, mainly because of the feasibility of approaching it with end-to-end neural models. In…

计算与语言 · 计算机科学 2023-03-14 Mohamed Nabih Ali , Alessio Brutti , Daniele Falavigna

Diffusion models have found great success in generating high quality, natural samples of speech, but their potential for density estimation for speech has so far remained largely unexplored. In this work, we leverage an unconditional…

音频与语音处理 · 电气工程与系统科学 2025-06-16 Danilo de Oliveira , Julius Richter , Jean-Marie Lemercier , Simon Welker , Timo Gerkmann

In this paper, we present a new objective prediction model for synthetic speech naturalness. It can be used to evaluate Text-To-Speech or Voice Conversion systems and works language independently. The model is trained end-to-end and based…

声音 · 计算机科学 2021-04-26 Gabriel Mittag , Sebastian Möller

Speaker recognition is a biometric modality that uses underlying speech information to determine the identity of the speaker. Speaker Identification (SID) under noisy conditions is one of the challenging topics in the field of speech…

声音 · 计算机科学 2019-08-02 Nursadul Mamun , Ria Ghosh , John H. L. Hansen

Objective: Currently, only behavioral speech understanding tests are available, which require active participation of the person being tested. As this is infeasible for certain populations, an objective measure of speech intelligibility is…

音频与语音处理 · 电气工程与系统科学 2021-11-30 Bernd Accou , Mohammad Jalilpour Monesi , Hugo Van hamme , Tom Francart

We propose a method using a long short-term memory (LSTM) network to estimate the noise power spectral density (PSD) of single-channel audio signals represented in the short time Fourier transform (STFT) domain. An LSTM network common to…

信号处理 · 电气工程与系统科学 2020-11-11 Xiaofei Li , Simon Leglaive , Laurent Girin , Radu Horaud

The perceptual task of speech quality assessment (SQA) is a challenging task for machines to do. Objective SQA methods that rely on the availability of the corresponding clean reference have been the primary go-to approaches for SQA.…

音频与语音处理 · 电气工程与系统科学 2021-10-19 Pranay Manocha , Buye Xu , Anurag Kumar

Uncertainty estimation for unlabeled data is crucial to active learning. With a deep neural network employed as the backbone model, the data selection process is highly challenging due to the potential over-confidence of the model…

机器学习 · 计算机科学 2024-02-14 Xingjian Li , Pengkun Yang , Yangcheng Gu , Xueying Zhan , Tianyang Wang , Min Xu , Chengzhong Xu