English
Related papers

Related papers: CasNet: Investigating Channel Robustness for Speec…

200 papers

Deep neural networks have become an indispensable technique for audio source separation (ASS). It was recently reported that a variant of CNN architecture called MMDenseNet was successfully employed to solve the ASS problem of estimating…

Sound · Computer Science 2018-05-30 Naoya Takahashi , Nabarun Goswami , Yuki Mitsufuji

In this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently with linear activation in the embedding layer. To understand…

Audio and Speech Processing · Electrical Eng. & Systems 2018-09-13 Suwon Shon , Hao Tang , James Glass

The presence of multiple talkers in the surrounding environment poses a difficult challenge for real-time speech communication systems considering the constraints on network size and complexity. In this paper, we present Personalized…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-09 Ritwik Giri , Shrikant Venkataramani , Jean-Marc Valin , Umut Isik , Arvindh Krishnaswamy

Recently, many deep learning based beamformers have been proposed for multi-channel speech separation. Nevertheless, most of them rely on extra cues known in advance, such as speaker feature, face image or directional information. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-08 Yanjie Fu , Haoran Yin , Meng Ge , Longbiao Wang , Gaoyan Zhang , Jianwu Dang , Chengyun Deng , Fei Wang

Deep gated convolutional networks have been proved to be very effective in single channel speech separation. However current state-of-the-art framework often considers training the gated convolutional networks in time-frequency (TF) domain.…

Sound · Computer Science 2019-03-19 Ziqiang Shi , Huibin Lin , Liu Liu , Rujie Liu , Shoji Hayakawa , Shouji Harada , Jiqing Han

The remarkable ability of humans to selectively focus on a target speaker in cocktail party scenarios is facilitated by binaural audio processing. In this paper, we present a binaural time-domain Target Speaker Extraction model based on the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-19 Hanyu Meng , Qiquan Zhang , Xiangyu Zhang , Vidhyasaharan Sethu , Eliathamby Ambikairajah

The performance of speech enhancement and separation systems in anechoic environments has been significantly advanced with the recent progress in end-to-end neural network architectures. However, the performance of such systems in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-17 Yi Luo , Cong Han , Nima Mesgarani

Achieving robust speech separation for overlapping speakers in various acoustic environments with noise and reverberation remains an open challenge. Although existing datasets are available to train separators for specific scenarios, they…

Sound · Computer Science 2024-08-30 Ke Chen , Jiaqi Su , Taylor Berg-Kirkpatrick , Shlomo Dubnov , Zeyu Jin

Neural multi-channel speech enhancement models, in particular those based on the U-Net architecture, demonstrate promising performance and generalization potential. These models typically encode input channels independently, and integrate…

Sound · Computer Science 2024-10-08 Ibrahim Aldarmaki , Thamar Solorio , Bhiksha Raj , Hanan Aldarmaki

To date, mainstream target speech separation (TSS) approaches are formulated to estimate the complex ratio mask (cRM) of the target speech in time-frequency domain under supervised deep learning framework. However, the existing deep models…

Sound · Computer Science 2021-09-08 Rongzhi Gu , Shi-Xiong Zhang , Yuexian Zou , Dong Yu

Recently, the ever-increasing capacity of large-scale annotated datasets has led to profound progress in stereo matching. However, most of these successes are limited to a specific dataset and cannot generalize well to other datasets. The…

Computer Vision and Pattern Recognition · Computer Science 2021-04-12 Zhelun Shen , Yuchao Dai , Zhibo Rao

Background: Active noise cancellation has been a subject of research for decades. Traditional techniques, like the Fast Fourier Transform, have limitations in certain scenarios. This research explores the use of deep neural networks (DNNs)…

Sound · Computer Science 2024-06-03 Brandon Colelough , Andrew Zheng

In daily listening environments, speech is always distorted by background noise, room reverberation and interference speakers. With the developing of deep learning approaches, much progress has been performed on monaural multi-speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Chao Ma , Dongmei Li , Xupeng Jia

The goal of this work is to develop a meeting transcription system that can recognize speech even when utterances of different speakers are overlapped. While speech overlaps have been regarded as a major obstacle in accurately transcribing…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-10 Takuya Yoshioka , Hakan Erdogan , Zhuo Chen , Xiong Xiao , Fil Alleva

End-to-end audio-conditioned latent diffusion models (LDMs) have been widely adopted for audio-driven portrait animation, demonstrating their effectiveness in generating lifelike and high-resolution talking videos. However, direct…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Chunyu Li , Chao Zhang , Weikai Xu , Jingyu Lin , Jinghui Xie , Weiguo Feng , Bingyue Peng , Cunjian Chen , Weiwei Xing

Target speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. We present IsoNet, a…

Sound · Computer Science 2026-05-18 Dinanath Padhya , Sajen Maharjan , Binita Adhikari , Ishwor Raj Pokharel

Speech separation is very important in real-world applications such as human-machine interaction, hearing aids devices, and automatic meeting transcription. In recent years, a significant improvement occurred towards the solution based on…

Sound · Computer Science 2024-08-29 Rawad Melhem , Assef Jafar , Oumayma Al Dakkak

In general, multi-channel source separation has utilized inter-microphone phase differences (IPDs) concatenated with magnitude information in time-frequency domain, or real and imaginary components stacked along the channel axis. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-01 Ui-Hyeop Shin , Bon Hyeok Ku , Hyung-Min Park

Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically rely on synthetic data pipelines, which may not reflect…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Runwu Shi , Kai Li , Chang Li , Jiang Wang , Sihan Tan , Kazuhiro Nakadai

Speaker Recognition is a challenging task with essential applications such as authentication, automation, and security. The SincNet is a new deep learning based model which has produced promising results to tackle the mentioned task. To…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-15 João Antônio Chagas Nunes , David Macêdo , Cleber Zanchettin