English
Related papers

Related papers: ShobdoSetu: A Data-Centric Framework for Bengali L…

200 papers

The prevalence of automatic speech recognition (ASR) systems in spoken language applications has increased significantly in recent years. Notably, many African languages lack sufficient linguistic resources to support the robustness of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-27 Amina Mardiyyah Rufai , Afolabi Abeeb , Esther Oduntan , Tayo Arulogun , Oluwabukola Adegboro , Daniel Ajisafe

Due to digitalization in everyday life, the need for automatically recognizing handwritten digits is increasing. Handwritten digit recognition is essential for numerous applications in various industries. Bengali ranks the fifth largest…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Muntarin Islam , Shabbir Ahmed Shuvo , Musarrat Saberin Nipun , Rejwan Bin Sulaiman , Jannatul Nayeem , Zubaer Haque , Md Mostak Shaikh , Md Sakib Ullah Sourav

Automatic speech recognition for low-resource languages remains fundamentally constrained by the scarcity of labeled data and computational resources required by state-of-the-art models. We present a systematic investigation into…

Computation and Language · Computer Science 2025-12-09 Srihari Bandarupalli , Bhavana Akkiraju , Charan Devarakonda , Vamsiraghusimha Narsinga , Anil Kumar Vuppala

Parkinson's disease (PD) poses a growing global health challenge, with Bangladesh experiencing a notable rise in PD-related mortality. Early detection of PD remains particularly challenging in resource-constrained settings, where…

Machine Learning · Computer Science 2025-05-20 Riad Hossain , Muhammad Ashad Kabir , Arat Ibne Golam Mowla , Animesh Chandra Roy , Ranjit Kumar Ghosh

This paper describes a new baseline system for automatic speech recognition (ASR) in the CHiME-4 challenge to promote the development of noisy ASR in speech processing communities by providing 1) state-of-the-art system with a simplified…

Sound · Computer Science 2018-03-28 Szu-Jui Chen , Aswin Shanmugam Subramanian , Hainan Xu , Shinji Watanabe

Automatic Speech Recognition (ASR) transcripts, especially in low-resource languages like Bangla, contain a critical ambiguity: word-word repetitions can be either Repetition Disfluency (unintentional ASR error/hesitation) or Morphological…

Computation and Language · Computer Science 2025-11-18 Zaara Zabeen Arpa , Sadnam Sakib Apurbo , Nazia Karim Khan Oishee , Ajwad Abrar

The Speaker Diarization and Recognition (SDR) task aims to predict "who spoke when and what" within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and dialogue systems.…

Sound · Computer Science 2026-01-06 Han Yin , Yafeng Chen , Chong Deng , Luyao Cheng , Hui Wang , Chao-Hong Tan , Qian Chen , Wen Wang , Xiangang Li

Sign Language Recognition (SLR) involves the automatic identification and classification of sign gestures from images or video, converting them into text or speech to improve accessibility for the hearing-impaired community. In Bangladesh,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Jubayer Ahmed Bhuiyan Shawon , Hasan Mahmud , Kamrul Hasan

Bengali social media platforms have witnessed a sharp increase in hate speech, disproportionately affecting women and adolescents. While datasets such as BD-SHS provide a basis for structured evaluation, most prior approaches rely on either…

Computation and Language · Computer Science 2026-02-24 Akif Islam , Mohd Ruhul Ameen

End-to-end speaker diarization enables accurate overlap-aware diarization by jointly estimating multiple speakers' speech activities in parallel. This approach is data-hungry, requiring a large amount of labeled conversational data, which…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-02 Shota Horiguchi , Atsushi Ando , Marc Delcroix , Naohiro Tawara

Large language models (LLMs) frequently exhibit performance biases against regional dialects of low-resource languages. However, frameworks to quantify these disparities remain scarce. We propose a two-phase framework to evaluate dialectal…

Computation and Language · Computer Science 2026-03-24 K. M. Jubair Sami , Dipto Sumit , Ariyan Hossain , Farig Sadeque

ASR has achieved remarkable global progress, yet African low-resource languages remain rigorously underrepresented, producing barriers to digital inclusion across the continent with more than +2000 languages. This systematic literature…

Recent improvements in multilingual ASR have not been equally distributed across languages and language varieties. To advance state-of-the-art (SOTA) ASR models, we present the Interspeech 2025 ML-SUPERB 2.0 Challenge. We construct a new…

Multi-speaker automatic speech recognition (ASR) aims to transcribe conversational speech involving multiple speakers, requiring the model to capture not only what was said, but also who said it and sometimes when it was spoken. Recent…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-27 Li Li , Ming Cheng , Weixin Zhu , Yannan Wang , Juan Liu , Ming Li

Speaker diarization is a fundamental task in speech processing that involves dividing an audio stream by speaker. Although state-of-the-art models have advanced performance in high-resource languages, low-resource languages such as Kurdish…

Diarization is a crucial component in meeting transcription systems to ease the challenges of speech enhancement and attribute the transcriptions to the correct speaker. Particularly in the presence of overlapping or noisy speech, these…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-06 Christoph Boeddeker , Tobias Cord-Landwehr , Reinhold Haeb-Umbach

Even state-of-the-art speaker diarization systems exhibit high variance in error rates across different datasets, representing numerous use cases and domains. Furthermore, comparing across systems requires careful application of best…

Sound · Computer Science 2025-08-07 Eduardo Pacheco , Atila Orhon , Berkin Durmus , Blaise Munyampirwa , Andrey Leonov

The evolving speech processing landscape is increasingly focused on complex scenarios like meetings or cocktail parties with multiple simultaneous speakers and far-field conditions. Existing methodologies for addressing these challenges…

Speaker diarization has been mainly developed based on the clustering of speaker embeddings. However, the clustering-based approach has two major problems; i.e., (i) it is not optimized to minimize diarization errors directly, and (ii) it…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-16 Yusuke Fujita , Naoyuki Kanda , Shota Horiguchi , Yawen Xue , Kenji Nagamatsu , Shinji Watanabe

Bangla is the seventh most spoken language by a total number of speakers in the world, and yet the development of an automated grammar checker in this language is an understudied problem. Bangla grammatical error detection is a task of…

Computation and Language · Computer Science 2024-11-14 Shayekh Bin Islam , Ridwanul Hasan Tanvir , Sihat Afnan
‹ Prev 1 4 5 6 7 8 10 Next ›