中文
相关论文

相关论文: ICASSP 2021 Acoustic Echo Cancellation Challenge: …

200 篇论文

A good joint training framework is very helpful to improve the performances of weakly supervised audio tagging (AT) and acoustic event detection (AED) simultaneously. In this study, we propose three methods to improve the best…

音频与语音处理 · 电气工程与系统科学 2022-02-15 Yunhao Liang , Yanhua Long , Yijie Li , Jiaen Liang , Yuping Wang

Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab…

This paper presents our work for the ICASSP 2026 Environmental Sound Deepfake Detection (ESDD) Challenge. The challenge is based on the large-scale EnvSDD dataset that consists of various synthetic environmental sounds. We focus on…

声音 · 计算机科学 2025-12-09 Candy Olivia Mawalim , Haotian Zhang , Shogo Okada

This paper addresses the challenge of speaker separation, which remains an active research topic despite the promising results achieved in recent years. These results, however, often degrade in real recording conditions due to the presence…

声音 · 计算机科学 2024-11-14 Rawad Melhem , Assef Jafar , Oumayma Al Dakkak

The growing prevalence of speech deepfakes has raised serious concerns, particularly in real-world scenarios such as telephone fraud and identity theft. While many anti-spoofing systems have demonstrated promising performance on…

音频与语音处理 · 电气工程与系统科学 2026-05-12 Tong Zhang , Yihuan Huang , Yanzhen Ren

Is pushing numbers on a single benchmark valuable in automatic speech recognition? Research results in acoustic modeling are typically evaluated based on performance on a single dataset. While the research community has coalesced around…

This thesis focuses on dealing with the task of acoustic scene classification (ASC), and then applied the techniques developed for ASC to a real-life application of detecting respiratory disease. To deal with ASC challenges, this thesis…

声音 · 计算机科学 2021-07-21 Lam Pham

Audio Packet Loss Concealment (PLC) is the hiding of gaps in audio streams caused by data transmission failures in packet switched networks. This is a common problem, and of increasing importance as end-to-end VoIP telephony and…

声音 · 计算机科学 2022-04-12 Lorenz Diener , Sten Sootla , Solomiya Branets , Ando Saabas , Robert Aichner , Ross Cutler

Deep neural networks (DNNs) have shown promising results for acoustic echo cancellation (AEC). But the DNN-based AEC models let through all near-end speakers including the interfering speech. In light of recent studies on personalized…

声音 · 计算机科学 2022-07-01 Shimin Zhang , Ziteng Wang , Yukai Ju , Yihui Fu , Yueyue Na , Qiang Fu , Lei Xie

The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge aims to benchmark and advance zero-shot spontaneous style voice cloning, particularly focusing on generating spontaneous behaviors in conversational speech. The challenge…

Speech enhancement is a task to improve the intelligibility and perceptual quality of degraded speech signal. Recently, neural networks based methods have been applied to speech enhancement. However, many neural network based methods…

声音 · 计算机科学 2021-02-22 Qiuqiang Kong , Haohe Liu , Xingjian Du , Li Chen , Rui Xia , Yuxuan Wang

While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities remains inadequate, partly due to limited public training…

Real-world speech communication is rarely affected by a single type of degradation. Instead, it suffers from a complex interplay of acoustic interference, codec compression, and, increasingly, secondary artifacts introduced by upstream…

声音 · 计算机科学 2025-12-30 Junan Zhang , Mengyao Zhu , Xin Xu , Hui Bu , Zhenhua Ling , Zhizheng Wu

ASR systems exhibit persistent performance disparities across accents, but whether these gaps reflect superficial biases or deep structural vulnerabilities remains unclear. We introduce ACES, a three-stage audit that extracts…

声音 · 计算机科学 2026-03-10 Swapnil Parekh

Single-channel speech enhancement approaches do not always improve automatic recognition rates in the presence of noise, because they can introduce distortions unhelpful for recognition. Following a trend towards end-to-end training of…

声音 · 计算机科学 2021-12-14 Peter Plantinga , Deblin Bagchi , Eric Fosler-Lussier

We present an iVector based Acoustic Scene Classification (ASC) system suited for real life settings where active foreground speech can be present. In the proposed system, each recording is represented by a fixed-length iVector that models…

音频与语音处理 · 电气工程与系统科学 2021-08-03 Siyuan Song , Brecht Desplanques , Celest De Moor , Kris Demuynck , Nilesh Madhu

Recently, more and more personalized speech enhancement systems (PSE) with excellent performance have been proposed. However, two critical issues still limit the performance and generalization ability of the model: 1) Acoustic environment…

音频与语音处理 · 电气工程与系统科学 2022-11-23 Xiaofeng Ge , Jiangyu Han , Haixin Guan , Yanhua Long

This challenge aims to evaluate the capabilities of audio encoders, especially in the context of multi-task learning and real-world applications. Participants are invited to submit pre-trained audio encoders that map raw waveforms to…

This technical report describes the IDLab submission for track 1 and 2 of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). This speaker verification competition focuses on short duration test recordings and cross-lingual trials.…

音频与语音处理 · 电气工程与系统科学 2021-09-10 Jenthe Thienpondt , Brecht Desplanques , Kris Demuynck

Error correction (EC) based on large language models is an emerging technology to enhance the performance of automatic speech recognition (ASR) systems. Generally, training data for EC are collected by automatically pairing a large set of…

计算与语言 · 计算机科学 2024-10-17 Takuma Udagawa , Masayuki Suzuki , Masayasu Muraoka , Gakuto Kurata