English
Related papers

Related papers: WTFormer: A Wavelet Conformer Network for MIMO Spe…

200 papers

The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder…

Sound · Computer Science 2017-03-16 Tsubasa Ochiai , Shinji Watanabe , Takaaki Hori , John R. Hershey

Supervised learning methods have shown effectiveness in estimating spatial acoustic parameters such as time difference of arrival, direct-to-reverberant ratio and reverberation time. However, they still suffer from the simulation-to-reality…

Sound · Computer Science 2024-09-10 Bing Yang , Xiaofei Li

Referring image segmentation aims to segment an object referred to by natural language expression from an image. The primary challenge lies in the efficient propagation of fine-grained semantic information from textual features to visual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Yichen Yan , Xingjian He , Sihan Chen , Jing Liu

This paper investigates robust semantic communications over multiple-input multiple-output (MIMO) fading channels. Current semantic communications over MIMO channels mainly focus on channel adaptive encoding and decoding, which lacks…

Information Theory · Computer Science 2024-07-09 Yiheng Duan , Tong Wu , Zhiyong Chen , Meixia Tao

Binaural speech enhancement faces a severe trade-off challenge, where state-of-the-art performance is achieved by computationally intensive architectures, while lightweight solutions often come at the cost of significant performance…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-26 Xikun Lu , Yujian Ma , Xianquan Jiang , Xuelong Wang , Jinqiu Sang

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both single-channel and multichannel conditions. However, severe performance degradation is still observed in the…

In this paper we provide three contributions to the field of channel sounding waveform design in asynchronous Multi-user (MU) MIMO systems. The first contribution is a derivation of the asynchronous MU-MIMO model and the conditions that the…

Information Theory · Computer Science 2013-02-20 Zhenhua Yu , Robert J. Baxley , Brett T. Walkenhorst , G. Tong Zhou

Massive multiple-input multiple-output (mMIMO) technology has transformed wireless communication by enhancing spectral efficiency and network capacity. This paper proposes a novel deep learning-based mMIMO precoder to tackle the complexity…

Signal Processing · Electrical Eng. & Systems 2025-02-14 Ali Hasanzadeh Karkan , Ahmed Ibrahim , Jean-François Frigon , François Leduc-Primeau

The performance of deep learning-based multi-channel speech enhancement methods often deteriorates when the geometric parameters of the microphone array change. Traditional approaches to mitigate this issue typically involve training on…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-03 Tianqin Zheng , Jilu Jin , Hanchen Pei , Gongping Huang , Jingdong Chen , Jacob Benesty

With the goal of improving spectral efficiency, complex rotation-based precoding and power allocation schemes are developed for two multiple-input multiple-output (MIMO) communication systems, namely, simultaneous wireless information and…

Information Theory · Computer Science 2021-11-30 Xinliang Zhang , Mojtaba Vaezi

Deep learning methods have brought substantial advancements in speech separation (SS). Nevertheless, it remains challenging to deploy deep-learning-based models on edge devices. Thus, identifying an effective way to compress these large…

Sound · Computer Science 2019-12-10 Chao-I Tuan , Yuan-Kuei Wu , Hung-yi Lee , Yu Tsao

Neural TTS has shown it can generate high quality synthesized speech. In this paper, we investigate the multi-speaker latent space to improve neural TTS for adapting the system to new speakers with only several minutes of speech or…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-04 Yan Deng , Lei He , Frank Soong

Downlink beamforming is an essential technology for wireless cellular networks; however, the design of beamforming vectors that maximize the weighted sum rate (WSR) is an NP-hard problem and iterative algorithms are typically applied to…

Signal Processing · Electrical Eng. & Systems 2022-06-28 Jingyuan Xia , Gunduz Deniz

We explore pretraining strategies including choice of base corpus with the aim of choosing the best strategy for zero-shot multi-speaker end-to-end synthesis. We also examine choice of neural vocoder for waveform synthesis, as well as…

Sound · Computer Science 2020-11-11 Erica Cooper , Xin Wang , Yi Zhao , Yusuke Yasuda , Junichi Yamagishi

Starting from first principles of wave propagation, we consider a multiple-input multiple-output (MIMO) representation of a communication system between two spatially-continuous volumes. This is the concept of holographic MIMO…

Signal Processing · Electrical Eng. & Systems 2022-09-30 Luca Sanguinetti , Antonio A. D'Amico , Merouane Debbah

Speech pre-processing techniques such as denoising, de-reverberation, and separation, are commonly employed as front-ends for various downstream speech processing tasks. However, these methods can sometimes be inadequate, resulting in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-17 Sirui Li , Shuai Wang , Zhijun Liu , Zhongjie Jiang , Yannan Wang , Haizhou Li

This paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recognition (ASR) backend. While the recently proposed Weighted…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-19 Zhaoheng Ni , Yong Xu , Meng Yu , Bo Wu , Shixiong Zhang , Dong Yu , Michael I Mandel

While deep learning-based models like transformers, have revolutionized time-series and vision tasks, they remain highly susceptible to noise and often overfit on noisy patterns rather than robust features. This issue is exacerbated in…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Ashish Bastola , Nishant Luitel , Hao Wang , Danda Pani Paudel , Roshani Poudel , Abolfazl Razi

The focus of this paper is on spatial precoding in correlated multi-antenna channels, where the number of independent data-streams is adapted to trade-off the data-rate with the transmitter complexity. Towards the goal of a low-complexity…

Information Theory · Computer Science 2009-09-29 Vasanthan Raghavan , Akbar Sayeed , Venu Veeravalli

Time series forecasting requires capturing patterns across multiple temporal scales while maintaining computational efficiency. This paper introduces AWGformer, a novel architecture that integrates adaptive wavelet decomposition with…

Machine Learning · Computer Science 2026-01-29 Wei Li
‹ Prev 1 4 5 6 7 8 10 Next ›