English
Related papers

Related papers: Direction-Preserving MIMO Speech Enhancement Using…

200 papers

Estimating a time-varying spatial covariance matrix for a beamforming algorithm is a challenging task, especially for wearable devices, as the algorithm must compensate for time-varying signal statistics due to rapid pose-changes. In this…

Sound · Computer Science 2021-12-10 Jonah Casebeer , Jacob Donley , Daniel Wong , Buye Xu , Anurag Kumar

In this paper, we introduce a neural network-based method for regional speech separation using a microphone array. This approach leverages novel spatial cues to extract the sound source not only from specified direction but also within…

Sound · Computer Science 2025-08-12 Yiheng Jiang , Haoxu Wang , Yafeng Chen , Gang Qiao , Biao Tian

This paper describes noisy speech recognition for an augmented reality headset that helps verbal communication within real multiparty conversational environments. A major approach that has actively been studied in simulated environments is…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-18 Yicheng Du , Aditya Arie Nugraha , Kouhei Sekiguchi , Yoshiaki Bando , Mathieu Fontaine , Kazuyoshi Yoshii

Orthogonal time frequency space (OTFS) modulation and massive multi-input multi-output (MIMO) are promising technologies for next generation wireless communication systems for their abilities to counteract the issue of high mobility with…

Signal Processing · Electrical Eng. & Systems 2025-04-14 Mingming Duan , Pengfei Zhang , Shun Zhang , Yao Ge , Octavia A. Dobre , Chau Yuen

Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech recognition capabilities. However, the ability of Speech LLMs to comprehend and process multi-channel audio with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-19 Jiamin Xie , Ju Lin , Yiteng Huang , Tyler Vuong , Zhaojiang Lin , Zhaojun Yang , Peng Su , Prashant Rawat , Sangeeta Srivastava , Ming Sun , Florian Metze

Hybrid precoding is an indispensable technique to harness the full potential of a multi-user massive multiple-input, multiple-output (MU-MMIMO) system. In this paper, we propose a new hybrid precoding approach that combines digital and…

Networking and Internet Architecture · Computer Science 2025-10-29 Azadeh Pourkabirian , Kai Li , Photios A. Stavrou , Wei Ni

Hybrid massive MIMO structures with lower hardware complexity and power consumption have been considered as a potential candidate for millimeter wave (mmWave) communications. Channel covariance information can be used for designing…

Signal Processing · Electrical Eng. & Systems 2019-06-25 Rui Hu , Jun Tong , Jiangtao Xi , Qinghua Guo , Yanguang Yu

In mmWave massive multiple-input multiple-output (mMIMO) systems, hybrid digital/analog beamforming has been recognized as an economic means to overcome the severe mmWave propagation loss. To facilitate beamforming for mmWace mMIMO, there…

Signal Processing · Electrical Eng. & Systems 2019-12-19 Shijian Gao , Xiang Cheng , Liuqing Yang

This paper proposes a dual-stage, low complexity, and reconfigurable technique to enhance the speech contaminated by various types of noise sources. Driven by input data and audio contents, the proposed dual-stage speech enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-18 Jun Yang , Nico Brailovsky

This work investigates the problem of spatial covariance matrix estimation in a millimeter-wave (mmWave) hybrid multiple-input multiple-output (MIMO) system with an emphasis on the basis-mismatch effect. The basis mismatch is prevalent in…

Signal Processing · Electrical Eng. & Systems 2019-12-11 Chethan Kumar Anjinappa , Ali Cafer Gurbuz , Yavuz Yapici , Ismail Guvenc

In this paper, we propose new accelerated update rules for rank-constrained spatial covariance model estimation, which efficiently extracts a directional target source in diffuse background noise.The naive updat e rule requires heavy…

Sound · Computer Science 2019-08-07 Yuki Kubo , Norihiro Takamune , Daichi Kitamura , Hiroshi Saruwatari

Uncertainty estimation in machine learning is paramount for enhancing the reliability and interpretability of predictive models, especially in high-stakes real-world scenarios. Despite the availability of numerous methods, they often pose a…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Anton Baumann , Thomas Roßberg , Michael Schmitt

This report focuses on algorithms that perform single-channel speech enhancement. The author of this report uses modulation-domain Kalman filtering algorithms for speech enhancement, i.e. noise suppression and dereverberation, in [1], [2],…

Sound · Computer Science 2018-11-02 Nikolaos Dionelis

This paper proposes a delayed subband LSTM network for online monaural (single-channel) speech enhancement. The proposed method is developed in the short time Fourier transform (STFT) domain. Online processing requires frame-by-frame signal…

Sound · Computer Science 2023-12-13 Xiaofei Li , Radu Horaud

We present a CNN architecture for speech enhancement from multichannel first-order Ambisonics mixtures. The data-dependent spatial filters, deduced from a mask-based approach, are used to help an automatic speech recognition engine to face…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-03 Amélie Bosca , Alexandre Guérin , Lauréline Perotin , Srđan Kitić

In this paper, we propose the coarse-to-fine optimization for the task of speech enhancement. Cosine similarity loss [1] has proven to be an effective metric to measure similarity of speech signals. However, due to the large variance of the…

Sound · Computer Science 2019-08-23 Jian Yao , Ahmad Al-Dahle

This paper proposes a channel estimation method for hybrid wideband multiple-input-multiple-output (MIMO) systems in high-frequency bands, including millimeter-wave (mmWave) and sub-terahertz (sub-THz), in the presence of beam squint…

Signal Processing · Electrical Eng. & Systems 2025-06-17 Kabuto Arai , Koji Ishibashi

This paper described the PCG-AIID system for L3DAS22 challenge in Task 1: 3D speech enhancement in office reverberant environment. We proposed a two-stage framework to address multi-channel speech denoising and dereverberation. In the first…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-22 Jingdong Li , Yuanyuan Zhu , Dawei Luo , Yun Liu , Guohui Cui , Zhaoxia Li

Keyword spotting (KWS) is crucial for many speech-driven applications, but robust KWS in noisy environments remains challenging. Conventional systems often rely on single-channel inputs and a cascaded pipeline separating front-end…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Rui Wang , Zhifei Zhang , Yu Gao , Xiaofeng Mou , Yi Xu

In hearing aid applications, an important objective is to accurately estimate the direction of arrival (DOA) of multiple speakers in noisy and reverberant environments. Recently, we proposed a binaural DOA estimation method, where the DOAs…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-11 Daniel Fejgin , Simon Doclo