English
Related papers

Related papers: Deep Learning Based Speech Beamforming

200 papers

The field of speech processing has undergone a transformative shift with the advent of deep learning. The use of multiple processing layers has enabled the creation of models capable of extracting intricate features from speech data. This…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-31 Ambuj Mehrish , Navonil Majumder , Rishabh Bhardwaj , Rada Mihalcea , Soujanya Poria

Beamforming is a signal processing technique. It has been studied in many areas such as radar, sonar, seismology and wireless communications, to name but a few. It can be used for a myriad of purposes, such as detecting the presence of a…

Other Computer Science · Computer Science 2012-12-27 Hidri Adel , Meddeb Souad , Abdulqadir Alaqeeli , Amiri Hamid

In real acoustic environment, speech enhancement is an arduous task to improve the quality and intelligibility of speech interfered by background noise and reverberation. Over the past years, deep learning has shown great potential on…

Sound · Computer Science 2021-05-07 Kanghao Zhang , Shulin He , Hao Li , Xueliang Zhang

In this paper, we explore an improved framework to train a monoaural neural enhancement model for robust speech recognition. The designed training framework extends the existing mixture invariant training criterion to exploit both unpaired…

Sound · Computer Science 2022-09-21 Jisi Zhang , Catalin Zorila , Rama Doddipatla , Jon Barker

This study presents a deep-learning framework for controlling multichannel acoustic feedback in audio devices. Traditional digital signal processing methods struggle with convergence when dealing with highly correlated noise such as…

Sound · Computer Science 2025-05-30 Yuan-Kuei Wu , Juan Azcarreta , Kashyap Patel , Buye Xu , Jung-Suk Lee , Sanha Lee , Ashutosh Pandey

It is highly desirable that speech enhancement algorithms can achieve good performance while keeping low latency for many applications, such as digital hearing aids, acoustically transparent hearing devices, and public address systems. To…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-01 Chengshi Zheng , Wenzhe Liu , Andong Li , Yuxuan Ke , Xiaodong Li

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

Computation and Language · Computer Science 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich

While models in audio and speech processing are becoming deeper and more end-to-end, they as a consequence need expensive training on large data, and are often brittle. We build on a classical model of human hearing and make it…

Sound · Computer Science 2024-09-16 Ruolan Leslie Famularo , Dmitry N. Zotkin , Shihab A. Shamma , Ramani Duraiswami

Thanks to advancements in deep learning, speech generation systems now power a variety of real-world applications, such as text-to-speech for individuals with speech disorders, voice chatbots in call centers, cross-linguistic speech…

We present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly. Given input audio containing speech corrupted by an additive background signal, the system aims to produce a processed…

Audio and Speech Processing · Electrical Eng. & Systems 2018-09-18 Francois G. Germain , Qifeng Chen , Vladlen Koltun

In ultrasound (US) imaging, various types of adaptive beamforming techniques have been investigated to improve the resolution and contrast-to-noise ratio of the delay and sum (DAS) beamformers. Unfortunately, the performance of these…

Image and Video Processing · Electrical Eng. & Systems 2020-02-25 Shujaat Khan , Jaeyoung Huh , Jong Chul Ye

Despite the remarkable progress recently made in distant speech recognition, state-of-the-art technology still suffers from a lack of robustness, especially when adverse acoustic conditions characterized by non-stationary noises and…

Computation and Language · Computer Science 2017-03-24 Mirco Ravanelli , Philemon Brakel , Maurizio Omologo , Yoshua Bengio

Speech recognition in adverse real-world environments is highly affected by reverberation and nonstationary background noise. A well-known strategy to reduce such undesired signal components in multi-microphone scenarios is spatial…

Sound · Computer Science 2017-08-08 Hendrik Barfuss , Christian Huemmer , Andreas Schwarz , Walter Kellermann

Speech clarity and spatial audio immersion are the two most critical factors in enhancing remote conferencing experiences. Existing methods are often limited: either due to the lack of spatial information when using only one microphone, or…

Sound · Computer Science 2025-07-14 Cheng Chi , Xiaoyu Li , Yuxuan Ke , Qunping Ni , Yao Ge , Xiaodong Li , Chengshi Zheng

This paper proposes a novel framework for audio deepfake detection with two main objectives: i) attaining the highest possible accuracy on available fake data, and ii) effectively performing continuous learning on new fake data in a…

Sound · Computer Science 2024-09-11 Tuan Duy Nguyen Le , Kah Kuan Teh , Huy Dat Tran

Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often yield suboptimal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Chihyun Liu , Jiaxuan Fan , Mingtung Sun , Michael Anthony , Mingsian R. Bai , Yu Tsao

The performance of speech enhancement algorithms in a multi-speaker scenario depends on correctly identifying the target speaker to be enhanced. Auditory attention decoding (AAD) methods allow to identify the target speaker which the…

Sound · Computer Science 2020-05-12 Ali Aroudi , Marc Delcroix , Tomohiro Nakatani , Keisuke Kinoshita , Shoko Araki , Simon Doclo

On-device directional hearing requires audio source separation from a given direction while achieving stringent human-imperceptible latency requirements. While neural nets can achieve significantly better performance than traditional…

Sound · Computer Science 2021-12-14 Anran Wang , Maruchi Kim , Hao Zhang , Shyamnath Gollakota

Multi-channel speech enhancement aims to recover clean speech from noisy multi-channel recordings. Most deep learning methods employ discriminative training, which can lead to non-linear distortions from regression-based objectives,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-26 Zhongweiyang Xu , Ashutosh Pandey , Juan Azcarreta , Zhaoheng Ni , Sanjeel Parekh , Buye Xu

Deep learning is an emerging technology that is considered one of the most promising directions for reaching higher levels of artificial intelligence. Among the other achievements, building computers that understand speech represents a…

Computation and Language · Computer Science 2017-12-19 Mirco Ravanelli