中文
相关论文

相关论文: A Multi-stage Low-latency Enhancement System for H…

200 篇论文

Although fully end-to-end speaker diarization systems have made significant progress in recent years, modular systems often achieve superior results in real-world scenarios due to their greater adaptability and robustness. Historically,…

音频与语音处理 · 电气工程与系统科学 2024-09-26 Ruoyu Wang , Shutong Niu , Gaobin Yang , Jun Du , Shuangqing Qian , Tian Gao , Jia Pan

To enhance the performance of end-to-end (E2E) speech recognition systems in noisy or low signal-to-noise ratio (SNR) conditions, this paper introduces NoisyD-CT, a novel tri-stage training framework built on the Conformer-Transducer…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Shuangyuan Chen , Shuang Wei , Dongxing Xu , Yanhua Long

Recent years have witnessed significant progress in multilingual automatic speech recognition (ASR), driven by the emergence of end-to-end (E2E) models and the scaling of multilingual datasets. Despite that, two main challenges persist in…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Zheshu Song , Jianheng Zhuo , Yifan Yang , Ziyang Ma , Shixiong Zhang , Xie Chen

The acoustic sensitivity of Autism Spectrum Disorder (ASD) individuals highly impacts their intelligibility in noisy urban environments. In this Letter, the disturbance sensing level is examined with perceptual listening tests that…

音频与语音处理 · 电气工程与系统科学 2025-05-16 Marcelo Pillonetto , Anderson Queiroz , Rosângela Coelho

This paper describes our NPU-Elevoc personalized speech enhancement system (NAPSE) for the 5th Deep Noise Suppression Challenge at ICASSP 2023. Based on the superior two-stage model TEA-PSE 2.0, our system particularly explores better…

音频与语音处理 · 电气工程与系统科学 2023-03-16 Xiaopeng Yan , Yindi Yang , Zhihao Guo , Liangliang Peng , Lei Xie

This paper introduces a parallel and asynchronous Transformer framework designed for efficient and accurate multilingual lip synchronization in real-time video conferencing systems. The proposed architecture integrates translation, speech…

多媒体 · 计算机科学 2025-12-23 Eren Caglar , Amirkia Rafiei Oskooei , Mehmet Kutanoglu , Mustafa Keles , Mehmet S. Aktas

We consider speech enhancement for signals picked up in one noisy environment that must be rendered to a listener in another noisy environment. For both far-end noise reduction and near-end listening enhancement, it has been shown that…

音频与语音处理 · 电气工程与系统科学 2024-02-06 Andreas J. Fuglsig , Jesper Jensen , Zheng-Hua Tan , Lars S. Bertelsen , Jens Christian Lindof , Jan Østergaard

Deploying speech enhancement (SE) systems in wearable devices, such as smart glasses, is challenging due to the limited computational resources on the device. Although deep learning methods have achieved high-quality results, their…

音频与语音处理 · 电气工程与系统科学 2025-08-21 Heitor R. Guimarães , Ke Tan , Juan Azcarreta , Jesus Alvarez , Prabhav Agrawal , Ashutosh Pandey , Buye Xu

Realizing edge intelligence consists of sensing, communication, training, and inference stages. Conventionally, the sensing and communication stages are executed sequentially, which results in excessive amount of dataset generation and…

信号处理 · 电气工程与系统科学 2022-01-25 Tong Zhang , Shuai Wang , Guoliang Li , Fan Liu , Guangxu Zhu , Rui Wang

This paper presents the architecture and performance of a novel Multilingual Automatic Speech Recognition (ASR) system developed by the Transsion Speech Team for Track 1 of the MLC-SLM 2025 Challenge. The proposed system comprises three key…

音频与语音处理 · 电气工程与系统科学 2025-08-22 Xiaoxiao Li , An Zhu , Youhai Jiang , Fengjie Zhu

We review current solutions and technical challenges for automatic speech recognition, keyword spotting, device arbitration, speech enhancement, and source localization in multidevice home environments to provide context for the INTERSPEECH…

音频与语音处理 · 电气工程与系统科学 2022-07-01 Gregory Ciccarelli , Jarred Barber , Arun Nair , Israel Cohen , Tao Zhang

Background noise and room reverberation are regarded as two major factors to degrade the subjective speech quality. In this paper, we propose an integrated framework to address simultaneous denoising and dereverberation under complicated…

声音 · 计算机科学 2021-06-25 Andong Li , Wenzhe Liu , Xiaoxue Luo , Guochen Yu , Chengshi Zheng , Xiaodong Li

Almost half a billion people world-wide suffer from disabling hearing loss. While hearing aids can partially compensate for this, a large proportion of users struggle to understand speech in situations with background noise. Here, we…

Integrated sensing and communication (ISAC), with sensing and communication sharing the same wireless resources and hardware, has the advantages of high spectrum efficiency and low hardware cost, which is regarded as one of the key…

信号处理 · 电气工程与系统科学 2023-06-09 Zhiqing Wei , Hanyang Qu , Wangjun Jiang , Kaifeng Han , Huici Wu , Zhiyong Feng

In this work, we propose a novel two-stage framework, called FaceShifter, for high fidelity and occlusion aware face swapping. Unlike many existing face swapping works that leverage only limited information from the target image when…

计算机视觉与模式识别 · 计算机科学 2020-09-16 Lingzhi Li , Jianmin Bao , Hao Yang , Dong Chen , Fang Wen

This paper describes the Royalflush speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription Challenge(M2MeT). Our system comprises speech enhancement, overlapped speech detection, speaker embedding…

声音 · 计算机科学 2022-02-21 Jingguang Tian , Xinhui Hu , Xinkang Xu

RADAR Challenge 2026 is an APSIPA Grand Challenge on Robust Audio Deepfake Recognition under Media Transformations, designed to simulate realistic media conditions in real-world audio distribution pipelines, including compression,…

音频与语音处理 · 电气工程与系统科学 2026-05-26 Hieu-Thi Luong , Xuechen Liu , Ivan Kukanov , Zheng Xin Chai , Kong Aik Lee

The introduction of AI and ML technologies into medical devices has revolutionized healthcare diagnostics and treatments. Medical device manufacturers are keen to maximize the advantages afforded by AI and ML by consolidating multiple…

软件工程 · 计算机科学 2024-02-08 Soham Sinha , Shekhar Dwivedi , Mahdi Azizian

Toward high-performance multilingual automatic speech recognition (ASR), various types of linguistic information and model design have demonstrated their effectiveness independently. They include language identity (LID), phoneme…

音频与语音处理 · 电气工程与系统科学 2025-01-10 Wei Liu , Jingyong Hou , Dong Yang , Muyong Cao , Tan Lee

This paper presents the Interspeech 2026 Audio Encoder Capability Challenge, a benchmark specifically designed to evaluate and advance the performance of pre-trained audio encoders as front-end modules for Large Audio Language Models…