English
Related papers

Related papers: DCHT: Deep Complex Hybrid Transformer for Speech E…

200 papers

This study presents a lightweight dual-domain super-resolution network (DDSRNet) that combines Spatial-Net with the discrete wavelet transform (DWT). Specifically, our proposed model comprises three main components: (1) a shallow feature…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Murat Karayaka , Usman Muhammad , Jorma Laaksonen , Md Ziaul Hoque , Tapio Seppänen

In recent years, a great deal of attention has been paid to the Transformer network for speech recognition tasks due to its excellent model performance. However, the Transformer network always involves heavy computation and large number of…

Sound · Computer Science 2023-04-12 Guangyong Wei , Zhikui Duan , Shiren Li , Guangguang Yang , Xinmei Yu , Junhua Li

In this paper we present a domain adaptation technique for formant estimation using a deep network. We first train a deep learning network on a small read speech dataset. We then freeze the parameters of the trained network and use several…

Computation and Language · Computer Science 2016-11-08 Yehoshua Dissen , Joseph Keshet , Jacob Goldberger , Cynthia Clopper

Deep Neural Networks (DNNs) often struggle to suppress noise at low signal-to-noise ratios (SNRs). This paper addresses speech enhancement in scenarios dominated by harmonic noise and proposes a framework that integrates…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-16 Giovanni Bologni , Nicolás Arrieta Larraza , Richard Heusdens , Richard C. Hendriks

Diffusion transformers have demonstrated remarkable generation quality, albeit requiring longer training iterations and numerous inference steps. In each denoising step, diffusion transformers encode the noisy inputs to extract the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Shuai Wang , Zhi Tian , Weilin Huang , Limin Wang

Curriculum learning begins to thrive in the speech enhancement area, which decouples the original spectrum estimation task into multiple easier sub-tasks to achieve better performance. Motivated by that, we propose a dual-branch…

Sound · Computer Science 2022-02-15 Guochen Yu , Andong Li , Chengshi Zheng , Yinuo Guo , Yutian Wang , Hui Wang

The widespread adoption of mobile communication technology has led to a severe shortage of spectrum resources, driving the development of cognitive radio technologies aimed at improving spectrum utilization, with spectrum sensing being the…

Signal Processing · Electrical Eng. & Systems 2025-04-11 Shilian Zheng , Zhihao Ye , Luxin Zhang , Keqiang Yue , Zhijin Zhao

Speech enhancement techniques based on deep learning have brought significant improvement on speech quality and intelligibility. Nevertheless, a large gain in speech quality measured by objective metrics, such as perceptual evaluation of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-06 Bo Wu , Meng Yu , Lianwu Chen , Yong Xu , Chao Weng , Dan Su , Dong Yu

Hand-crafted spatial features, such as inter-channel intensity difference (IID) and inter-channel phase difference (IPD), play a fundamental role in recent deep learning based dual-microphone speech enhancement (DMSE) systems. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-04 Xinmeng Xu , Rongzhi Gu , Yuexian Zou

Transformer has shown advanced performance in speech separation, benefiting from its ability to capture global features. However, capturing local features and channel information of audio sequences in speech separation is equally important.…

Sound · Computer Science 2023-03-08 Zhaoxi Mu , Xinyu Yang , Wenjing Zhu

In multichannel speech enhancement, both spectral and spatial information are vital for discriminating between speech and noise. How to fully exploit these two types of information and their temporal dynamics remains an interesting research…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-17 Yujie Yang , Changsheng Quan , Xiaofei Li

This paper proposes a model that integrates sub-band processing and deep filtering to fully exploit information from the target time-frequency (TF) bin and its surrounding TF bins for single-channel speech enhancement. The sub-band module…

Sound · Computer Science 2025-06-03 Shenghui Lu , Hukai Huang , Jinanglong Yao , Kaidi Wang , Qingyang Hong , Lin Li

Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would…

Computation and Language · Computer Science 2018-02-16 Kaizhi Qian , Yang Zhang , Shiyu Chang , Xuesong Yang , Dinei Florencio , Mark Hasegawa-Johnson

In this work, we present CleanUNet 2, a speech denoising model that combines the advantages of waveform denoiser and spectrogram denoiser and achieves the best of both worlds. CleanUNet 2 uses a two-stage framework inspired by popular…

Machine Learning · Computer Science 2023-09-13 Zhifeng Kong , Wei Ping , Ambrish Dantrey , Bryan Catanzaro

Spoken language models (SLMs) have gained increasing attention with advancements in text-based, decoder-only language models. SLMs process text and speech, enabling simultaneous speech understanding and generation. This paper presents…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-01 Heng-Jui Chang , Hongyu Gong , Changhan Wang , James Glass , Yu-An Chung

In this contribution, we investigate the effectiveness of deep fusion of text and audio features for categorical and dimensional speech emotion recognition (SER). We propose a novel, multistage fusion method where the two information…

Machine Learning · Computer Science 2023-03-27 Andreas Triantafyllopoulos , Uwe Reichel , Shuo Liu , Stephan Huber , Florian Eyben , Björn W. Schuller

The deep learning-based speech enhancement (SE) methods always take the clean speech's waveform or time-frequency spectrum feature as the learning target, and train the deep neural network (DNN) by reducing the error loss between the DNN's…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-02 Yuewei Zhang , Huanbin Zou , Jie Zhu

Electrocardiography (ECG) signals can be considered as multi-variable time-series. The state-of-the-art ECG data classification approaches, based on either feature engineering or deep learning techniques, treat separately spectral and time…

Machine Learning · Computer Science 2023-11-10 Che Liu , Sibo Cheng , Weiping Ding , Rossella Arcucci

This paper introduces a deep neural network model for subband-based speech synthesizer. The model benefits from the short bandwidth of the subband signals to reduce the complexity of the time-domain speech generator. We employed the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-28 Azam Rabiee , Geonmin Kim , Tae-Ho Kim , Soo-Young Lee

Dialect variation hampers automatic recognition of bird calls collected by passive acoustic monitoring. We address the problem on DB3V, a three-region, ten-species corpus of 8-s clips, and propose a deployable framework built on Time-Delay…

Sound · Computer Science 2025-09-29 Jiani Ding , Qiyang Sun , Alican Akman , Björn W. Schuller