English
Related papers

Related papers: UX-NET: Filter-and-Process-based Improved U-Net fo…

200 papers

This work introduces UDPNet, a novel architecture designed to accelerate the reverse diffusion process in speech synthesis. Unlike traditional diffusion models that rely on timestep embeddings and shared network parameters, UDPNet unrolls…

Sound · Computer Science 2025-06-12 Peter Ochieng

Recent research has made significant progress in designing fusion modules for audio-visual speech separation. However, they predominantly focus on multi-modal fusion at a single temporal scale of auditory and visual features without…

Sound · Computer Science 2024-02-05 Kai Li , Runxuan Yang , Fuchun Sun , Xiaolin Hu

We introduce CrossNet, a complex spectral mapping approach to speaker separation and enhancement in reverberant and noisy conditions. The proposed architecture comprises an encoder layer, a global multi-head self-attention module, a…

Sound · Computer Science 2024-03-07 Vahid Ahmadi Kalkhorani , DeLiang Wang

Semi-supervised learning is increasingly popular in medical image segmentation due to its ability to leverage large amounts of unlabeled data to extract additional information. However, most existing semi-supervised segmentation methods…

Computer Vision and Pattern Recognition · Computer Science 2024-08-19 Rong Wu , Dehua Li , Cong Zhang

Recent studies have increasingly acknowledged the advantages of incorporating visual data into speech enhancement (SE) systems. In this paper, we introduce a novel audio-visual SE approach, termed DCUC-Net (deep complex U-Net with conformer…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-10 Shafique Ahmed , Chia-Wei Chen , Wenze Ren , Chin-Jou Li , Ernie Chu , Jun-Cheng Chen , Amir Hussain , Hsin-Min Wang , Yu Tsao , Jen-Cheng Hou

This paper introduces an improved target speaker extractor, referred to as Speakerfilter-Pro, based on our previous Speakerfilter model. The Speakerfilter uses a bi-direction gated recurrent unit (BGRU) module to characterize the target…

Sound · Computer Science 2020-10-27 Shulin He , Hao Li , Xueliang Zhang

This research paper presents a novel deep learning-based neural network architecture, named Y-Net, for achieving music source separation. The proposed architecture performs end-to-end hybrid source separation by extracting features from…

We investigate the viability of a variational U-Net architecture for denoising of single-channel audio data. Deep network speech enhancement systems commonly aim to estimate filter masks, or opt to work on the waveform signal, potentially…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-04 Eike J. Nustede , Jörn Anemüller

Image segmentation is a fundamental task in image analysis and clinical practice. The current state-of-the-art techniques are based on U-shape type encoder-decoder networks with skip connections, called U-Net. Despite the powerful…

Computer Vision and Pattern Recognition · Computer Science 2023-02-02 Chun-Wun Cheng , Christina Runkel , Lihao Liu , Raymond H Chan , Carola-Bibiane Schönlieb , Angelica I Aviles-Rivero

We address monaural multi-speaker-image separation in reverberant conditions, aiming at separating mixed speakers but preserving the reverberation of each speaker. A straightforward approach for this task is to directly train end-to-end DNN…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Jingqi Sun , Shulin He , Ruizhe Pang , Zhong-Qiu Wang

We propose TF-GridNet for speech separation. The model is a novel deep neural network (DNN) integrating full- and sub-band modeling in the time-frequency (T-F) domain. It stacks several blocks, each consisting of an intra-frame full-band…

In this paper, we propose TitaNet, a novel neural network architecture for extracting speaker representations. We employ 1D depth-wise separable convolutions with Squeeze-and-Excitation (SE) layers with global context followed by channel…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-12 Nithin Rao Koluguri , Taejin Park , Boris Ginsburg

The successful deployment of deep learning-based acoustic echo and noise reduction (AENR) methods in consumer devices has spurred interest in developing low-complexity solutions, while emphasizing the need for robust performance in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-05 Shrishti Saha Shetu , Naveen Kumar Desiraju , Wolfgang Mack , Emanuël A. P. Habets

This paper proposes a neural network based speech separation method using spatially distributed microphones. Unlike with traditional microphone array settings, neither the number of microphones nor their spatial arrangement is known in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Dongmei Wang , Zhuo Chen , Takuya Yoshioka

In this paper, we describe the work that we have done to participate in Task1 of the ConferencingSpeech2021 challenge. This task set a goal to develop the solution for multi-channel speech enhancement in a real-time manner. We propose a…

Signal Processing · Electrical Eng. & Systems 2021-04-06 Vasiliy Kuzmin , Fyodor Kravchenko , Artem Sokolov , Jie Geng

The advent of deep learning has led to the prevalence of deep neural network architectures for monaural music source separation, with end-to-end approaches that operate directly on the waveform level increasingly receiving research…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-09 Christos Garoufis , Athanasia Zlatintsi , Petros Maragos

Segmenting ultrasound images is critical for various medical applications, but it offers significant challenges due to ultrasound images' inherent noise and unpredictability. To address these challenges, we proposed EUIS-Net, a CNN network…

Image and Video Processing · Electrical Eng. & Systems 2024-08-23 Shahzaib Iqbal , Hasnat Ahmed , Muhammad Sharif , Madiha Hena , Tariq M. Khan , Imran Razzak

Multi-speaker speech recognition has been one of the keychallenges in conversation transcription as it breaks the singleactive speaker assumption employed by most state-of-the-artspeech recognition systems. Speech separation is consideredas…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-08 Jian Wu , Zhuo Chen , Jinyu Li , Takuya Yoshioka , Zhili Tan , Ed Lin , Yi Luo , Lei Xie

Recently, deep learning-based beamforming algorithms have shown promising performance in target speech extraction tasks. However, most systems do not fully utilize spatial information. In this paper, we propose a target speech extraction…

Sound · Computer Science 2023-06-29 Aoqi Guo , Junnan Wu , Peng Gao , Wenbo Zhu , Qinwen Guo , Dazhi Gao , Yujun Wang

Speech separation has been extensively studied to deal with the cocktail party problem in recent years. All related approaches can be divided into two categories: time-frequency domain methods and time domain methods. In addition, some…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-31 Fan-Lin Wang , Yu-Huai Peng , Hung-Shin Lee , Hsin-Min Wang