Bengali-Loop: Community Benchmarks for Long-Form Bangla ASR and Speaker Diarization
Abstract
Bengali (Bangla) remains under-resourced in long-form speech technology despite its wide use. We present Bengali-Loop, two community benchmarks to address this gap: (1) a long-form ASR corpus of 191 recordings (158.6 hours, 792k words) from 11 YouTube channels, collected via a reproducible subtitle-extraction pipeline and human-in-the-loop transcript verification; and (2) a speaker diarization corpus of 24 recordings (22 hours, 5,744 annotated segments) with fully manual speaker-turn labels in CSV format. Both benchmarks target realistic multi-speaker, long-duration content (e.g., Bangla drama/natok). We establish baselines (Tugstugi: 34.07% WER; pyannote.audio: 40.08% DER) and provide standardized evaluation protocols (WER/CER, DER), annotation rules, and data formats to support reproducible benchmarking and future model development for Bangla long-form ASR and diarization.
Keywords
Cite
@article{arxiv.2602.14291,
title = {Bengali-Loop: Community Benchmarks for Long-Form Bangla ASR and Speaker Diarization},
author = {H. M. Shadman Tabib and Istiak Ahmmed Rifti and Abdullah Muhammed Amimul Ehsan and Somik Dasgupta and Md Zim Mim Siddiqee Sowdha and Abrar Jahin Sarker and Md. Rafiul Islam Nijamy and Tanvir Hossain and Mst. Metaly Khatun and Munzer Mahmood and Rakesh Debnath and Gourab Biswas and Asif Karim and Wahid Al Azad Navid and Masnoon Muztahid and Fuad Ahmed Udoy and Shahad Shahriar Rahman and Md. Tashdiqur Rahman Shifat and Most. Sonia Khatun and Mushfiqur Rahman and Md. Miraj Hasan and Anik Saha and Mohammad Ninad Mahmud Nobo and Soumik Bhattacharjee and Tusher Bhomik and Ahmmad Nur Swapnil and Shahriar Kabir},
journal= {arXiv preprint arXiv:2602.14291},
year = {2026}
}