English
Related papers

Related papers: Time-domain Ad-hoc Array Speech Enhancement Using …

200 papers

This paper describes noisy speech recognition for an augmented reality headset that helps verbal communication within real multiparty conversational environments. A major approach that has actively been studied in simulated environments is…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-18 Yicheng Du , Aditya Arie Nugraha , Kouhei Sekiguchi , Yoshiaki Bando , Mathieu Fontaine , Kazuyoshi Yoshii

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. To improve robustness of speaker recognition system performance in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Yanpei Shi , Qiang Huang , Thomas Hain

In this paper, we address the problem of multichannel speech enhancement in the short-time Fourier transform (STFT) domain. A long short-time memory (LSTM) network takes as input a sequence of STFT coefficients associated with a frequency…

Sound · Computer Science 2020-09-24 Xiaofei LI , Radu Horaud

This paper makes several contributions to automatic lyrics transcription (ALT) research. Our main contribution is a novel variant of the Multistreaming Time-Delay Neural Network (MTDNN) architecture, called MSTRE-Net, which processes the…

Sound · Computer Science 2021-08-06 Emir Demirel , Sven Ahlbäck , Simon Dixon

The dominant speech separation models are based on complex recurrent or convolution neural network that model speech sequences indirectly conditioning on context, such as passing information through many intermediate states in recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-17 Jingjing Chen , Qirong Mao , Dong Liu

Nowadays, one practical limitation of deep neural network (DNN) is its high degree of specialization to a single task or domain (e.g., one visual domain). It motivates researchers to develop algorithms that can adapt DNN model to multiple…

Computer Vision and Pattern Recognition · Computer Science 2021-10-07 Li Yang , Adnan Siraj Rakin , Deliang Fan

A promising approach for multi-microphone speech separation involves two deep neural networks (DNN), where the predicted target speech from the first DNN is used to compute signal statistics for time-invariant minimum variance…

Sound · Computer Science 2021-10-04 Zhong-Qiu Wang , Gordon Wichern , Jonathan Le Roux

Ambisonics encoding of microphone array signals can enable various spatial audio applications, such as virtual reality or telepresence, but it is typically designed for uniformly-spaced spherical microphone arrays. This paper proposes a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-12 Mikko Heikkinen , Archontis Politis , Tuomas Virtanen

A neural network is essentially a high-dimensional complex mapping model by adjusting network weights for feature fitting. However, the spectral bias in network training leads to unbearable training epochs for fitting the high-frequency…

Signal Processing · Electrical Eng. & Systems 2021-06-22 Zhi Zeng , Pengpeng Shi , Fulei Ma , Peihan Qi

Deep learning-based speech enhancement has shown unprecedented performance in recent years. The most popular mono speech enhancement frameworks are end-to-end networks mapping the noisy mixture into an estimate of the clean speech. With…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-02 Bahareh Tolooshams , Kazuhito Koishida

Despite the remarkable progress recently made in distant speech recognition, state-of-the-art technology still suffers from a lack of robustness, especially when adverse acoustic conditions characterized by non-stationary noises and…

Computation and Language · Computer Science 2017-03-24 Mirco Ravanelli , Philemon Brakel , Maurizio Omologo , Yoshua Bengio

In this work, we investigate the generalization of a multi-channel learning-based replay speech detector, which employs adaptive beamforming and detection, across different microphone arrays. In general, deep neural network-based microphone…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-09 Michael Neri , Tuomas Virtanen

Almost in every heavily computation-dependent application, from 6G communication systems to autonomous driving platforms, a large portion of computing should be near to the client side. Edge computing (AI at Edge) in mobile devices is one…

Hardware Architecture · Computer Science 2024-07-29 Seyed Nima Omidsajedi , Rekha Reddy , Jianming Yi , Jan Herbst , Christoph Lipps , Hans Dieter Schotten

Multi-channel speech enhancement aims to recover clean speech from noisy multi-channel recordings. Most deep learning methods employ discriminative training, which can lead to non-linear distortions from regression-based objectives,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-26 Zhongweiyang Xu , Ashutosh Pandey , Juan Azcarreta , Zhaoheng Ni , Sanjeel Parekh , Buye Xu

Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network based masking…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-18 Shubo Lv , Yihui Fu , Yukai Jv , Lei Xie , Weixin Zhu , Wei Rao , Yannan Wang

A promising approach for steering auditory attention in complex listening environments relies on Auditory Attention Decoding (AAD), which aim to identify the attended speech stream in a multiple speaker scenario from neural recordings.…

This study proposes a trainable adaptive window switching (AWS) method and apply it to a deep-neural-network (DNN) for speech enhancement in the modified discrete cosine transform domain. Time-frequency (T-F) mask processing in the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-02-21 Yuma Koizumi , Noboru Harada , Yoichi Haneda

Extracting direct-path spatial feature is crucial for sound source localization in adverse acoustic environments. This paper proposes the IPDnet, a neural network that estimates direct-path inter-channel phase difference (DP-IPD) of sound…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-14 Yabo Wang , Bing Yang , Xiaofei Li

Financial time-series forecasting is one of the most challenging domains in the field of time-series analysis. This is mostly due to the highly non-stationary and noisy nature of financial time-series data. With progressive efforts of the…

Machine Learning · Computer Science 2022-01-17 Mostafa Shabani , Dat Thanh Tran , Martin Magris , Juho Kanniainen , Alexandros Iosifidis

Spiking Neural Networks (SNNs) present a more energy-efficient alternative to Artificial Neural Networks (ANNs) by harnessing spatio-temporal dynamics and event-driven spikes. Effective utilization of temporal information is crucial for…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Minje Kim , Minjun Kim , Xu Yang