English
Related papers

Related papers: Domain Adaptation of the Pyannote Diarization Pipe…

200 papers

Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like…

Computation and Language · Computer Science 2026-01-15 Fadhil Muhammad , Alwin Djuliansah , Adrian Aryaputra Hamzah , Kurniawati Azizah

Speaker diarization is the task of partitioning audio into segments according to speaker identity, answering the question of "who spoke when" in multi-speaker conversation recordings. While diarization is an essential task for many…

Sound · Computer Science 2025-10-01 Luca A. Lanzendörfer , Florian Grötschla , Cesare Blaser , Roger Wattenhofer

This study examines domain effects in speaker diarization for African-accented English. We evaluate multiple production and open systems on general and clinical dialogues under a strict DER protocol that scores overlap. A consistent domain…

Computation and Language · Computer Science 2025-09-29 Chibuzor Okocha , Kelechi Ezema , Christan Grant

We introduce DAS (Domain Adaptation with Synthetic data), a novel domain adaptation framework for pre-trained ASR model, designed to efficiently adapt to various language-defined domains without requiring any real data. In particular, DAS…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-23 Minh Tran , Yutong Pang , Debjyoti Paul , Laxmi Pandey , Kevin Jiang , Jinxi Guo , Ke Li , Shun Zhang , Xuedong Zhang , Xin Lei

An ideal speech recognition model has the capability to transcribe speech accurately under various characteristics of speech signals, such as speaking style (read and spontaneous), speech context (formal and informal), and background noise…

Computation and Language · Computer Science 2024-10-15 Aulia Adila , Dessi Lestari , Ayu Purwarianti , Dipta Tanaya , Kurniawati Azizah , Sakriani Sakti

Automatic Speech Recognition (ASR) and speaker diarization in Bangla remain challenging due to long form recordings, diverse acoustic conditions, and significant speaker variability. This work addresses these two core tasks in Bangla spoken…

Speech recognizers trained on close-talking speech do not generalize to distant speech and the word error rate degradation can be as large as 40% absolute. Most studies focus on tackling distant speech recognition as a separate problem,…

Computation and Language · Computer Science 2018-06-14 Hao Tang , Wei-Ning Hsu , Francois Grondin , James Glass

In this work, we propose a method to create domain-sensitive speech recognition models that utilize textual domain information by conditioning its generation on a given text prompt. This is accomplished by fine-tuning a pre-trained,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-09 Feng-Ting Liao , Yung-Chieh Chan , Yi-Chang Chen , Chan-Jan Hsu , Da-shan Shiu

This report presents the system developed by the ABSP Laboratory team for the third DIHARD speech diarization challenge. Our main contribution in this work is to develop a simple and efficient solution for acoustic domain dependent speech…

Sound · Computer Science 2021-01-26 A Kishore Kumar , Shefali Waldekar , Goutam Saha , Md Sahidullah

Diarization partitions an audio stream into segments based on the voices of the speakers. Real-time diarization systems that include an enrollment step should limit enrollment training samples to reduce user interaction time. Although…

Sound · Computer Science 2022-08-09 Dirk Padfield , Daniel J. Liebling

This study focuses on the development of Indonesian Automatic Speech Recognition (ASR) using the XLSR-53 pre-trained model, the XLSR stands for cross-lingual speech representations. The use of this XLSR-53 pre-trained model is to…

Computation and Language · Computer Science 2023-08-23 Panji Arisaputra , Amalia Zahra

Speaker adaptation techniques provide a powerful solution to customise automatic speech recognition (ASR) systems for individual users. Practical application of unsupervised model-based speaker adaptation techniques to data intensive…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-16 Jiajun Deng , Xurong Xie , Tianzi Wang , Mingyu Cui , Boyang Xue , Zengrui Jin , Guinan Li , Shujie Hu , Xunying Liu

Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies ``who spoke when'', is an essential component of the automated analysis. However, publicly available…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-13 Anfeng Xu , Tiantian Feng , Helen Tager-Flusberg , Catherine Lord , Shrikanth Narayanan

Named entity recognition (NER) is an important task in NLP, which is all the more challenging in conversational domain with their noisy facets. Moreover, conversational texts are often available in limited amount, making supervised tasks…

Computation and Language · Computer Science 2019-02-22 Rezka Leonandya , Fariz Ikhwantri

This study investigates the impact of integrating a dataset of disordered speech recordings ($\sim$1,000 hours) into the fine-tuning of a near state-of-the-art ASR baseline system. Contrary to what one might expect, despite the data being…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-22 Jimmy Tobin , Katrin Tomanek , Subhashini Venugopalan

This study investigates the performance of personalized automatic speech recognition (ASR) for recognizing disordered speech using small amounts of per-speaker adaptation data. We trained personalized models for 195 individuals with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-12 Jimmy Tobin , Katrin Tomanek

Domain adaptation is often hampered by exceedingly small target datasets and inaccessible source data. These conditions are prevalent in speech verification, where privacy policies and/or languages with scarce speech resources limit the…

Sound · Computer Science 2024-06-11 Shlomo Salo Elia , Aviad Malachi , Vered Aharonson , Gadi Pinkas

End-to-end neural diarization has evolved considerably over the past few years, but data scarcity is still a major obstacle for further improvements. Self-supervised learning methods such as WavLM have shown promising performance on several…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-22 Jiangyu Han , Federico Landini , Johan Rohdin , Anna Silnova , Mireia Diez , Lukas Burget

We describe our end-to-end system for Bengali long-form speech recognition (ASR) and speaker diarization submitted to the DL Sprint 4.0 competition on Kaggle. Bengali presents substantial challenges for both tasks: a large phoneme…

Computation and Language · Computer Science 2026-02-26 MD. Sagor Chowdhury , Adiba Fairooz Chowdhury

In this paper, we investigate the transferability of pre-trained language models to low-resource Indonesian local languages through the task of sentiment analysis. We evaluate both zero-shot performance and adapter-based transfer on ten…

Computation and Language · Computer Science 2025-07-03 Rifki Afina Putri
‹ Prev 1 2 3 10 Next ›