English
Related papers

Related papers: M-CIF: Multi-Scale Alignment For CIF-Based Non-Aut…

200 papers

Children's automatic speech recognition (ASR) is always difficult due to, in part, the data scarcity problem, especially for kindergarten-aged kids. When data are scarce, the model might overfit to the training data, and hence good starting…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-28 Yunzheng Zhu , Ruchao Fan , Abeer Alwan

Recent advances in large language models have facilitated the development of unified speech language models (SLMs) capable of supporting multiple speech tasks within a shared architecture. However, tasks such as automatic speech recognition…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Yuke Si , Runyan Yang , Yingying Gao , Junlan Feng , Chao Deng , Shilei Zhang

Automatic Speech Recognition (ASR) has shown remarkable progress, yet it still faces challenges in real-world distant scenarios across various array topologies each with multiple recording devices. The focal point of the CHiME-7 Distant ASR…

Sound · Computer Science 2023-12-18 Bingshen Mu , Pengcheng Guo , Dake Guo , Pan Zhou , Wei Chen , Lei Xie

In this paper, we present our overall efforts to improve the performance of a code-switching speech recognition system using semi-supervised training methods from lexicon learning to acoustic modeling, on the South East Asian…

Computation and Language · Computer Science 2018-06-19 Pengcheng Guo , Haihua Xu , Lei Xie , Eng Siong Chng

We propose a novel one-pass multiple ASR systems joint compression and quantization approach using an all-in-one neural model. A single compression cycle allows multiple nested systems with varying Encoder depths, widths, and quantization…

It has been shown that the intelligibility of noisy speech can be improved by speech enhancement (SE) algorithms. However, monaural SE has not been established as an effective frontend for automatic speech recognition (ASR) in noisy…

Sound · Computer Science 2024-03-12 Yufeng Yang , Ashutosh Pandey , DeLiang Wang

In the field of artificial intelligence, self-supervised learning has demonstrated superior generalization capabilities by leveraging large-scale unlabeled datasets for pretraining, which is especially critical for wireless communication…

Signal Processing · Electrical Eng. & Systems 2025-03-04 Jun Jiang , Wenjun Yu , Yunfan Li , Yuan Gao , Shugong Xu

Despite notable advancements in automatic speech recognition (ASR), performance tends to degrade when faced with adverse conditions. Generative error correction (GER) leverages the exceptional text comprehension capabilities of large…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-08 Bingshen Mu , Yangze Li , Qijie Shao , Kun Wei , Xucheng Wan , Naijun Zheng , Huan Zhou , Lei Xie

The existing audio datasets are predominantly tailored towards single languages, overlooking the complex linguistic behaviors of multilingual communities that engage in code-switching. This practice, where individuals frequently mix two or…

Sound · Computer Science 2025-03-04 Peng Xie , Kani Chen

Thanks to the rise of self-supervised learning, automatic speech recognition (ASR) systems now achieve near-human performance on a wide variety of datasets. However, they still lack generalization capability and are not robust to domain…

Machine Learning · Computer Science 2023-03-15 Lucas Maison , Yannick Estève

Automatic Cued Speech Recognition (ACSR) provides an intelligent human-machine interface for visual communications, where the Cued Speech (CS) system utilizes lip movements and hand gestures to code spoken language for hearing-impaired…

Computer Vision and Pattern Recognition · Computer Science 2023-02-28 Lei Liu , Li Liu

Training large foundation models using self-supervised objectives on unlabeled data, followed by fine-tuning on downstream tasks, has emerged as a standard procedure. Unfortunately, the efficacy of this approach is often constrained by both…

The rapid development of 5G New Radio (NR) and millimeter-wave (mmWave) communication systems highlights the critical importance of maintaining accurate phase synchronization to ensure reliable and efficient communication. This study…

Signal Processing · Electrical Eng. & Systems 2024-12-10 Desire Guel , Flavien Herve Somda , Boureima Zerbo , Oumarou Sie

In this paper, we describe recent performance improvements to the production Marchex speech recognition system for our spontaneous customer-to-business telephone conversations. In our previous work, we focused on in-domain language and…

Computation and Language · Computer Science 2019-05-03 Seongjun Hahm , Iroro Orife , Shane Walker , Jason Flaks

In speech enhancement (SE), phase estimation is important for perceptual quality, so many methods take clean speech's complex short-time Fourier transform (STFT) spectrum or the complex ideal ratio mask (cIRM) as the learning target. To…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-12 Yuewei Zhang , Huanbin Zou , Jie Zhu

The success of the multilingual automatic speech recognition systems empowered many voice-driven applications. However, measuring the performance of such systems remains a major challenge, due to its dependency on manually transcribed…

Computation and Language · Computer Science 2023-04-04 Shammur Absar Chowdhury , Ahmed Ali

To advance integrated sensing and communications (ISAC) in sixth-generation (6G) extremely large-scale multiple-input multiple-output (XL-MIMO) networks, a low-complexity compressed sensing (CS)-based dictionary design is proposed for…

Signal Processing · Electrical Eng. & Systems 2026-03-31 Ruiyun Zhang , Zhaolin Wang , Zhiqing Wei , Yuanwei Liu , Zehui Xiong , Zhiyong Feng

Integrate-and-Fire Time Encoding Machine (IF-TEM) is a power-efficient asynchronous sampler that converts analog signals into non-uniform time-domain samples. Adaptive IF-TEM (AIF-TEM) improves this machine by adapting its process to the…

Signal Processing · Electrical Eng. & Systems 2025-11-05 Vered Karp , Aseel Omar , Alejandro Cohen

Fine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We propose an incremental semi-supervised learning pipeline…

Recent advancements have showcased the potential of handheld millimeter-wave (mmWave) imaging, which applies synthetic aperture radar (SAR) principles in portable settings. However, existing studies addressing handheld motion errors either…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Yadong Li , Dongheng Zhang , Ruixu Geng , Jincheng Wu , Yang Hu , Qibin Sun , Yan Chen