English
Related papers

Related papers: Multi-channel Opus compression for far-field autom…

200 papers

This paper presents a novel donor data selection method to enhance low-resource automatic speech recognition (ASR). While ASR performs well in high-resource languages, its accuracy declines in low-resource settings due to limited training…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-30 Shunsuke Mitsumori , Sara Kashiwagi , Keitaro Tanaka , Shigeo Morishima

Audiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-world scenarios across various domains due to noisy acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Yihan Wu , Yichen Lu , Yifan Peng , Xihua Wang , Ruihua Song , Shinji Watanabe

In this paper, we analyzed how audio-visual speech enhancement can help to perform the ASR task in a cocktail party scenario. Therefore we considered two simple end-to-end LSTM-based models that perform single-channel audio-visual speech…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-28 Luca Pasa , Giovanni Morrone , Leonardo Badino

Massive MIMO is a promising technique for future 5G communications due to its high spectrum and energy efficiency. To realize its potential performance gain, accurate channel estimation is essential. However, due to massive number of…

Information Theory · Computer Science 2016-11-17 Zhen Gao , Linglong Dai , Wei Dai , Byonghyo Shim , Zhaocheng Wang

Mobile and embedded machine learning developers frequently have to compromise between two inferior on-device deployment strategies: sacrifice accuracy and aggressively shrink their models to run on dedicated low-power cores; or sacrifice…

Machine Learning · Computer Science 2023-03-17 Haiguang Li , Trausti Thormundsson , Ivan Poupyrev , Nicholas Gillian

Compared with automatic speech recognition (ASR), the human auditory system is more adept at handling noise-adverse situations, including environmental noise and channel distortion. To mimic this adeptness, auditory models have been widely…

Computation and Language · Computer Science 2016-09-16 Peng Dai , Xue Teng , Frank Rudzicz , Ing Yann Soon

Online Transformer-based automatic speech recognition (ASR) systems have been extensively studied due to the increasing demand for streaming applications. Recently proposed Decoder-end Adaptive Computation Steps (DACS) algorithm for online…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-27 Mohan Li , Catalin Zorila , Rama Doddipatla

A data compression system capable of providing real-time streaming of high-resolution continuous point-on-wave (CPOW) and phasor measurement unit (PMU) measurements is proposed. Referred to as adaptive subband compression (ASBC), the…

Signal Processing · Electrical Eng. & Systems 2024-05-01 Xinyi Wang , Yilu Liu , Lang Tong

Personalized speech enhancement has been a field of active research for suppression of speechlike interferers such as competing speakers or TV dialogues. Compared with single channel approaches, multichannel PSE systems can be more…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-17 Yicheng Hsu , Yonghan Lee , Mingsian R. Bai

Voice communication using the air conduction microphone in noisy environments suffers from the degradation of speech audibility. Bone conduction microphones (BCM) are robust against ambient noises but suffer from limited effective bandwidth…

Sound · Computer Science 2021-12-28 Yuang Li , Yuntao Wang , Xin Liu , Yuanchun Shi , Shao-fu Shih

Single-word Automatic Speech Recognition (ASR) is a challenging task due to the lack of linguistic context and sensitivity to noise, pronunciation variation, and channel artifacts, especially in low-resource, communication-critical domains…

Sound · Computer Science 2026-01-30 Manali Sharma , Riya Naik , Buvaneshwari G

In this paper, we present UR-AIR system submission to the logical access (LA) and the speech deepfake (DF) tracks of the ASVspoof 2021 Challenge. The LA and DF tasks focus on synthetic speech detection (SSD), i.e. detecting text-to-speech…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-05 Xinhui Chen , You Zhang , Ge Zhu , Zhiyao Duan

Pre-trained models, especially self-supervised learning (SSL) models, have demonstrated impressive results in automatic speech recognition (ASR) task. While most applications of SSL models focus on leveraging continuous representations as…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Zehan Li , Yan Yang , Xueqing Li , Jian Kang , Xiao-Lei Zhang , Jie Li

There is a growing interest in signaling schemes that operate in the wideband regime due to the crowded frequency spectrum. However, a downside of the wideband regime is that obtaining channel state information is costly, and the capacity…

Signal Processing · Electrical Eng. & Systems 2022-06-01 Kathleen Yang , Diana C. Gonzalez , Yonina C. Eldar , Muriel Medard

This paper investigates over-the-air computation (AirComp) in the context of multiple-access time-varying multipath channels. We focus on a scenario where devices with high mobility transmit their sensing data to a fusion center (FC) for…

Information Theory · Computer Science 2024-03-19 Xinyu Huang , Henrik Hellstrom , Carlo Fischione

We study the problem of uplink compression for cell-free multi-input multi-output networks with limited fronthaul capacity. In compress-forward mode, remote radio heads (RRHs) compress the received signal and forward it to a central unit…

Information Theory · Computer Science 2025-06-16 Zehua Li , Jingjie Wei , Raviraj Adve

The orthogonal frequency division multiplexing is a very efficient modulation technique that can achieve very high throughput by transmitting many carriers simultaneously and it is spectrally efficient because of the proximity of the…

Signal Processing · Electrical Eng. & Systems 2022-06-07 Usha S. M , Mahesh H. B

While Separate Source-Channel Coding (SSCC) retains the practical benefits of modular system design, its effectiveness in noisy text transmission is fundamentally constrained by the fragility of autoregressive source decoding. In low-SNR…

Information Theory · Computer Science 2026-05-08 Ziqiong Wang , Rongpeng Li

The high-intensity, repetitive noise associated with functional magnetic resonance imaging hinders on-line monitoring of subjects' speech and/or recording speech signals suitable for off-line analysis. The proposed algorithm enhances the…

Sound · Computer Science 2012-07-26 Satrajit S. Ghosh

Recently, speech recognition with ad-hoc microphone arrays has received much attention. It is known that channel selection is an important problem of ad-hoc microphone arrays, however, this topic seems far from explored in speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-15 Junqi Chen , Xiao-Lei Zhang