中文
相关论文

相关论文: AMNet: An Acoustic Model Network for Enhanced Mand…

200 篇论文

We investigate a novel cross-lingual multi-speaker text-to-speech synthesis approach for generating high-quality native or accented speech for native/foreign seen/unseen speakers in English and Mandarin. The system consists of three…

音频与语音处理 · 电气工程与系统科学 2019-11-27 Zhaoyu Liu , Brian Mak

This paper presents a novel deep learning architecture for acoustic model in the context of Automatic Speech Recognition (ASR), termed as MixNet. Besides the conventional layers, such as fully connected layers in DNN-HMM and memory cells in…

音频与语音处理 · 电气工程与系统科学 2021-12-03 Vishwanath Pratap Singh , Shakti P. Rath , Abhishek Pandey

In speech enhancement, the lack of clear structural characteristics in the target speech phase requires the use of conservative and cumbersome network frameworks. It seems difficult to achieve competitive performance using direct methods…

声音 · 计算机科学 2023-06-08 Liang Liu , Haixin Guan , Jinlong Ma , Wei Dai , Guangyong Wang , Shaowei Ding

While traditional statistical signal processing model-based methods can derive the optimal estimators relying on specific statistical assumptions, current learning-based methods further promote the performance upper bound via deep neural…

声音 · 计算机科学 2022-03-17 Andong Li , Chengshi Zheng , Ziyang Zhang , Xiaodong Li

The human auditory system has the ability to selectively focus on key speech elements in an audio stream while giving secondary attention to less relevant areas such as noise or distortion within the background, dynamically adjusting its…

音频与语音处理 · 电气工程与系统科学 2026-04-09 Nursadul Mamun , John H. L. Hansen

Beamforming has been extensively investigated for multi-channel audio processing tasks. Recently, learning-based beamforming methods, sometimes called \textit{neural beamformers}, have achieved significant improvements in both signal…

音频与语音处理 · 电气工程与系统科学 2019-10-02 Yi Luo , Enea Ceolini , Cong Han , Shih-Chii Liu , Nima Mesgarani

Recent years have witnessed significant improvement in ASR systems to recognize spoken utterances. However, it is still a challenging task for noisy and out-of-domain data, where substitution and deletion errors are prevalent in the…

音频与语音处理 · 电气工程与系统科学 2021-06-17 Mukuntha Narayanan Sundararaman , Ayush Kumar , Jithendra Vepa

The front-end module in multi-channel automatic speech recognition (ASR) systems mainly use microphone array techniques to produce enhanced signals in noisy conditions with reverberation and echos. Recently, neural network (NN) based…

声音 · 计算机科学 2020-11-19 Yuxiang Kong , Jian Wu , Quandong Wang , Peng Gao , Weiji Zhuang , Yujun Wang , Lei Xie

Compensation for channel mismatch and noise interference is essential for robust automatic speech recognition. Enhanced speech has been introduced into the multi-condition training of acoustic models to improve their generalization ability.…

声音 · 计算机科学 2022-11-24 Hung-Shin Lee , Pin-Yuan Chen , Yao-Fei Cheng , Yu Tsao , Hsin-Min Wang

FullSubNet is our recently proposed real-time single-channel speech enhancement network that achieves outstanding performance on the Deep Noise Suppression (DNS) Challenge dataset. A number of variants of FullSubNet have been proposed, but…

音频与语音处理 · 电气工程与系统科学 2023-03-08 Xiang Hao , Xiaofei Li

Complex-valued processing has brought deep learning-based speech enhancement and signal extraction to a new level. Typically, the process is based on a time-frequency (TF) mask which is applied to a noisy spectrogram, while complex masks…

音频与语音处理 · 电气工程与系统科学 2022-02-02 Hendrik Schröter , Alberto N. Escalante-B. , Tobias Rosenkranz , Andreas Maier

Due to the over-fitting problem caused by imbalance samples, there is still room to improve the performance of data-driven automatic modulation classification (AMC) in noisy scenarios. By fully considering the signal characteristics, an AMC…

信号处理 · 电气工程与系统科学 2022-03-08 Hao Shi , Qi Peng , Yiqi Zhuang

Mispronunciation Detection and Diagnosis (MDD) systems, leveraging Automatic Speech Recognition (ASR), face two main challenges in Mandarin Chinese: 1) The two-stage models create an information gap between the phoneme or tone…

声音 · 计算机科学 2024-06-10 Xintong Wang , Mingqian Shi , Ye Wang

Neural network applications generally benefit from larger-sized models, but for current speech enhancement models, larger scale networks often suffer from decreased robustness to the variety of real-world use cases beyond what is…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Umut Isik , Ritwik Giri , Neerad Phansalkar , Jean-Marc Valin , Karim Helwani , Arvindh Krishnaswamy

This paper makes several contributions to automatic lyrics transcription (ALT) research. Our main contribution is a novel variant of the Multistreaming Time-Delay Neural Network (MTDNN) architecture, called MSTRE-Net, which processes the…

声音 · 计算机科学 2021-08-06 Emir Demirel , Sven Ahlbäck , Simon Dixon

Neural speech synthesis models can synthesize high quality speech but typically require a high computational complexity to do so. In previous work, we introduced LPCNet, which uses linear prediction to significantly reduce the complexity of…

音频与语音处理 · 电气工程与系统科学 2022-02-24 Jean-Marc Valin , Umut Isik , Paris Smaragdis , Arvindh Krishnaswamy

We present aTENNuate, a simple deep state-space autoencoder configured for efficient online raw speech enhancement in an end-to-end fashion. The network's performance is primarily evaluated on raw speech denoising, with additional…

声音 · 计算机科学 2025-06-17 Yan Ru Pei , Ritik Shrivastava , FNU Sidharth

In this paper, we propose a speaker verification method by an Attentive Multi-scale Convolutional Recurrent Network (AMCRN). The proposed AMCRN can acquire both local spatial information and global sequential information from the input…

音频与语音处理 · 电气工程与系统科学 2023-06-02 Yanxiong Li , Zhongjie Jiang , Wenchang Cao , Qisheng Huang

Two factors have proven to be very important to the performance of semantic segmentation models: global context and multi-level semantics. However, generating features that capture both factors always leads to high computational complexity,…

计算机视觉与模式识别 · 计算机科学 2021-03-11 Qi Song , Kangfu Mei , Rui Huang

Recently, there has been an increasing interest in neural speech synthesis. While the deep neural network achieves the state-of-the-art result in text-to-speech (TTS) tasks, how to generate a more emotional and more expressive speech is…

计算与语言 · 计算机科学 2021-06-24 Chenye Cui , Yi Ren , Jinglin Liu , Feiyang Chen , Rongjie Huang , Ming Lei , Zhou Zhao