中文
相关论文

相关论文: SICRN: Advancing Speech Enhancement through State …

200 篇论文

In this paper, we summarize recent progresses made in deep learning based acoustic models and the motivation and insights behind the surveyed techniques. We first discuss acoustic models that can effectively exploit variable-length…

音频与语音处理 · 电气工程与系统科学 2018-04-30 Dong Yu , Jinyu Li

Environmental Sound Classification (ESC) is an important and challenging problem, and feature representation is a critical and even decisive factor in ESC. Feature representation ability directly affects the accuracy of sound…

声音 · 计算机科学 2019-08-19 Tianhao Qiao , Shunqing Zhang , Zhichao Zhang , Shan Cao , Shugong Xu

This paper introduces a practical approach for leveraging a real-time deep learning model to alternate between speech enhancement and joint speech enhancement and separation depending on whether the input mixture contains one or two active…

音频与语音处理 · 电气工程与系统科学 2023-10-17 Kashyap Patel , Anton Kovalyov , Issa Panahi

Sound Event Localization and Detection refers to the problem of identifying the presence of independent or temporally-overlapped sound sources, correctly identifying to which sound class it belongs, estimating their spatial directions while…

音频与语音处理 · 电气工程与系统科学 2024-01-02 Francesca Ronchini , Daniel Arteaga , Andrés Pérez-López

Teleconferencing is becoming essential during the COVID-19 pandemic. However, in real-world applications, speech quality can deteriorate due to, for example, background interference, noise, or reverberation. To solve this problem, target…

音频与语音处理 · 电气工程与系统科学 2022-05-02 Yicheng Hsu , Yonghan Lee , Mingsian R. Bai

In real-time speech recognition applications, the latency is an important issue. We have developed a character-level incremental speech recognition (ISR) system that responds quickly even during the speech, where the hypotheses are…

计算与语言 · 计算机科学 2016-06-29 Kyuyeon Hwang , Wonyong Sung

Convolutional neural networks (CNN) are one of the best-performing neural network architectures for environmental sound classification (ESC). Recently, temporal attention mechanisms have been used in CNN to capture the useful information…

声音 · 计算机科学 2020-05-22 Helin Wang , Yuexian Zou , Dading Chong , Wenwu Wang

Encoder transformer models compress information from all tokens in a sequence into a single [CLS] token to represent global context. This approach risks diluting fine-grained or hierarchical features, leading to information loss in…

计算与语言 · 计算机科学 2025-09-23 Asif Shahriar , Rifat Shahriyar , M Saifur Rahman

Recently, single-image super-resolution has made great progress owing to the development of deep convolutional neural networks (CNNs). The vast majority of CNN-based models use a pre-defined upsampling operator, such as bicubic…

计算机视觉与模式识别 · 计算机科学 2019-08-28 Xin Yang , Haiyang Mei , Jiqing Zhang , Ke Xu , Baocai Yin , Qiang Zhang , Xiaopeng Wei

Speech Emotion Recognition (SER) systems often degrade in performance when exposed to the unpredictable acoustic interference found in real-world environments. Additionally, the opacity of deep learning models hinders their adoption in…

声音 · 计算机科学 2025-12-23 Sudip Chakrabarty , Pappu Bishwas , Rajdeep Chatterjee

This paper proposes a framework for modeling sound change that combines deep learning and iterative learning. Acquisition and transmission of speech is modeled by training generations of Generative Adversarial Networks (GANs) on unannotated…

计算与语言 · 计算机科学 2021-09-23 Gašper Beguš

In this paper, we propose a speaker verification method by an Attentive Multi-scale Convolutional Recurrent Network (AMCRN). The proposed AMCRN can acquire both local spatial information and global sequential information from the input…

音频与语音处理 · 电气工程与系统科学 2023-06-02 Yanxiong Li , Zhongjie Jiang , Wenchang Cao , Qisheng Huang

Speech enhancement techniques based on deep learning have brought significant improvement on speech quality and intelligibility. Nevertheless, a large gain in speech quality measured by objective metrics, such as perceptual evaluation of…

音频与语音处理 · 电气工程与系统科学 2020-07-06 Bo Wu , Meng Yu , Lianwu Chen , Yong Xu , Chao Weng , Dan Su , Dong Yu

Speech separation models are used for isolating individual speakers in many speech processing applications. Deep learning models have been shown to lead to state-of-the-art (SOTA) results on a number of speech separation benchmarks. One…

声音 · 计算机科学 2023-03-13 William Ravenscroft , Stefan Goetze , Thomas Hain

Noise suppression in seismic data processing is a crucial research focus for enhancing subsequent imaging and reservoir prediction. Deep learning has shown promise in computer vision and holds significant potential for seismic data…

地球物理 · 物理学 2024-08-06 Fei Li , Zhenbin Xia , Dawei Liu , Xiaokai Wang , Wenchao Chen , Juan Chen , Leiming Xu

Classification of EEG signals using shallow Convolutional Neural Networks (CNNs) is a prevalent and successful approach across a variety of fields. Most of these models use independent one-dimensional (1D) convolutional layers along the…

机器学习 · 计算机科学 2026-05-06 Laurits Dixen , Stefan Heinrich , Paolo Burelli

Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier transform(STFT)…

声音 · 计算机科学 2019-01-03 Xingjian Du , Mengyao Zhu , Xuan Shi , Xinpeng Zhang , Wen Zhang , Jingdong Chen

Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or…

声音 · 计算机科学 2025-01-20 Shengkui Zhao , Zexu Pan , Kun Zhou , Yukun Ma , Chong Zhang , Bin Ma

Citrinet is an end-to-end convolutional Connectionist Temporal Classification (CTC) based automatic speech recognition (ASR) model. To capture local and global contextual information, 1D time-channel separable convolutions combined with…

计算与语言 · 计算机科学 2022-09-02 Xianchao Wu

The deep learning based time-domain models, e.g. Conv-TasNet, have shown great potential in both single-channel and multi-channel speech enhancement. However, many experiments on the time-domain speech enhancement model are done in…

音频与语音处理 · 电气工程与系统科学 2021-10-28 Wangyou Zhang , Jing Shi , Chenda Li , Shinji Watanabe , Yanmin Qian
‹ 上一页 1 8 9 10 下一页 ›