English
Related papers

Related papers: The CHiME-8 DASR Challenge for Generalizable and A…

200 papers

In this paper, we present the speaker diarization system for the Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) from team DKU_DukeECE. As the highly overlapped speech exists in the dataset, we employ an x-vector-based…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-08 Weiqing Wang , Xiaoyi Qin , Ming Li

Noisy situations cause huge problems for suffers of hearing loss as hearing aids often make the signal more audible but do not always restore the intelligibility. In noisy settings, humans routinely exploit the audio-visual (AV) nature of…

Sound · Computer Science 2019-09-24 Mandar Gogate , Kia Dashtipour , Ahsan Adeel , Amir Hussain

Identifying the identity of the speaker of short segments in human dialogue has been considered one of the most challenging problems in speech signal processing. Speaker representations of short speech segments tend to be unreliable,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-23 Tae Jin Park , Manoj Kumar , Shrikanth Narayanan

This paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage both audio and video information to improve the…

The deep learning based time-domain models, e.g. Conv-TasNet, have shown great potential in both single-channel and multi-channel speech enhancement. However, many experiments on the time-domain speech enhancement model are done in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-28 Wangyou Zhang , Jing Shi , Chenda Li , Shinji Watanabe , Yanmin Qian

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020. We open…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-28 Chandan K A Reddy , Harishchandra Dubey , Vishak Gopal , Ross Cutler , Sebastian Braun , Hannes Gamper , Robert Aichner , Sriram Srinivasan

This paper presents the TEA-ASLP's system submitted to the MLC-SLM 2025 Challenge, addressing multilingual conversational automatic speech recognition (ASR) in Task I and speech diarization ASR in Task II. For Task I, we enhance Ideal-LLM…

Sound · Computer Science 2025-07-25 Hongfei Xue , Kaixun Huang , Zhikai Zhou , Shen Huang , Shidong Shang

Long-context understanding is crucial for many NLP applications, yet transformers struggle with efficiency due to the quadratic complexity of self-attention. Sparse attention methods alleviate this cost but often impose static, predefined…

Computation and Language · Computer Science 2025-06-16 Hanzhi Zhang , Heng Fan , Kewei Sha , Yan Huang , Yunhe Feng

In this report, we describe the speaker diarization (SD) and language diarization (LD) systems developed by our team for the Second DISPLACE Challenge, 2024. Our contributions were dedicated to Track 1 for SD and Track 2 for LD in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-25 Nikhil Raghav , Subhajit Saha , Md Sahidullah , Swagatam Das

This paper describes the DKU-MSXF submission to track 4 of the VoxCeleb Speaker Recognition Challenge 2023 (VoxSRC-23). Our system pipeline contains voice activity detection, clustering-based diarization, overlapped speech detection, and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-21 Ming Cheng , Weiqing Wang , Xiaoyi Qin , Yuke Lin , Ning Jiang , Guoqing Zhao , Ming Li

Despite the strong modeling power of neural network acoustic models, speech enhancement has been shown to deliver additional word error rate improvements if multi-channel data is available. However, there has been a longstanding debate…

Computation and Language · Computer Science 2019-09-27 Catalin Zorila , Christoph Boeddeker , Rama Doddipatla , Reinhold Haeb-Umbach

Denoising language models (DLMs) have been proposed as a powerful alternative to traditional language models (LMs) for automatic speech recognition (ASR), motivated by their ability to use bidirectional context and adapt to a specific ASR…

Neural and Evolutionary Computing · Computer Science 2025-12-16 Dorian Koch , Albert Zeyer , Nick Rossenbach , Ralf Schlüter , Hermann Ney

This paper describes the XMUSPEECH speaker recognition and diarisation systems for the VoxCeleb Speaker Recognition Challenge 2021. For track 2, we evaluate two systems including ResNet34-SE and ECAPA-TDNN. For track 4, an important part of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-07 Jie Wang , Fuchuang Tong , Zhicong Chen , Lin Li , Qingyang Hong , Haodong Zhou

In this paper, we study several microphone channel selection and weighting methods for robust automatic speech recognition (ASR) in noisy conditions. For channel selection, we investigate two methods based on the maximum likelihood (ML)…

Sound · Computer Science 2016-10-04 Zhaofeng Zhang , Xiong Xiao , Longbiao Wang , EngSiong Chng , Haizhou Li

Dysarthric speech recognition faces challenges from severity variations and disparities relative to normal speech. Conventional approaches individually fine-tune ASR models pre-trained on normal speech per patient to prevent feature…

Sound · Computer Science 2025-08-27 Qing Xiao , Yingshan Peng , PeiPei Zhang

This report describes the submission of the DKU-DukeECE-Lenovo team to the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2021 track 4. Our system including a voice activity detection (VAD) model, a speaker embedding model, two…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-08 Weiqing Wang , Danwei Cai , Qingjian Lin , Lin Yang , Junjie Wang , Jin Wang , Ming Li

This paper describes the BUCEA speaker diarization system for the 2022 VoxCeleb Speaker Recognition Challenge. Voxsrc-22 provides the development set and test set of VoxConverse, and we mainly use the test set of VoxConverse for parameter…

Sound · Computer Science 2022-09-21 Ruohua Zhou , Yuxuan Du , Chenlei Hu

This paper presents the details of the SRIB-LEAP submission to the ConferencingSpeech challenge 2021. The challenge involved the task of multi-channel speech enhancement to improve the quality of far field speech from microphone arrays in a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-25 R G Prithvi Raj , Rohit Kumar , M K Jayesh , Anurenjan Purushothaman , Sriram Ganapathy , M A Basha Shaik

We present the task description of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2023 Challenge Task 2: ``First-shot unsupervised anomalous sound detection (ASD) for machine condition monitoring''. The main goal is…

Personalizing Automatic Speech Recognition (ASR) for non-normative speech remains challenging because data collection is labor-intensive and model training is technically complex. To address these limitations, we propose Adapt4Me, a…

Human-Computer Interaction · Computer Science 2026-03-23 Niclas Pokel , Yiming Zhao , Pehuén Moure , Yingqiang Gao , Roman Böhringer
‹ Prev 1 8 9 10 Next ›