English
Related papers

Related papers: A Crowdsourced Open-Source Kazakh Speech Corpus an…

200 papers

For many of the 700 million illiterate people around the world, speech recognition technology could provide a bridge to valuable information and services. Yet, those most in need of this technology are often the most underserved by it. In…

Machine Learning · Computer Science 2021-04-28 Moussa Doumbouya , Lisa Einstein , Chris Piech

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce…

Computation and Language · Computer Science 2025-09-23 Yuhang Dai , Ziyu Zhang , Shuai Wang , Longhao Li , Zhao Guo , Tianlun Zuo , Shuiyuan Wang , Hongfei Xue , Chengyou Wang , Qing Wang , Xin Xu , Hui Bu , Jie Li , Jian Kang , Binbin Zhang , Lei Xie

Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt dataset for safety evaluation across eleven categories covering common risk areas such as…

Computation and Language · Computer Science 2026-05-29 Wajdi Zaghouani , Shimaa Amer Ibrahim , Aruzhan Muratbek , Olzhasbek Zhakenov , Adiya Akhmetzhanova

Whisper and other large-scale automatic speech recognition models have made significant progress in performance. However, their performance on many low-resource languages, such as Kazakh, is not satisfactory. It is worth researching how to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-08 Jinpeng Li , Yu Pu , Qi Sun , Wei-Qiang Zhang

Parliamentary transcripts provide a valuable resource to understand the reality and know about the most important facts that occur over time in our societies. Furthermore, the political debates captured in these transcripts facilitate…

Recent advances in audio-language models have demonstrated remarkable success on short, segment-level speech tasks. However, real-world applications such as meeting transcription, spoken document understanding, and conversational analysis…

We introduce VoxPopuli, a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages. It is the largest open data to date for unsupervised representation learning as well as semi-supervised learning.…

Computation and Language · Computer Science 2021-07-28 Changhan Wang , Morgane Rivière , Ann Lee , Anne Wu , Chaitanya Talnikar , Daniel Haziza , Mary Williamson , Juan Pino , Emmanuel Dupoux

In this paper, we present a new Russian and Kazakh database (with about 95% of Russian and 5% of Kazakh words/sentences respectively) for offline handwriting recognition. A few pre-processing and segmentation procedures have been developed…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Daniyar Nurseitov , Kairat Bostanbekov , Daniyar Kurmankhojayev , Anel Alimova , Abdelrahman Abdallah

Dubbed series are gaining a lot of popularity in recent years with strong support from major media service providers. Such popularity is fueled by studies that showed that dubbed versions of TV shows are more popular than their subtitled…

Computation and Language · Computer Science 2022-03-08 Massa Baali , Wassim El-Hajj , Ahmed Ali

Slovak remains a low-resource language for automatic speech recognition (ASR), with fewer than 100 hours of publicly available training data. We present SloPal, a comprehensive Slovak parliamentary corpus comprising 330,000…

Computation and Language · Computer Science 2026-03-17 Erik Božík , Marek Šuppa

We introduce HarperValleyBank, a free, public domain spoken dialog corpus. The data simulate simple consumer banking interactions, containing about 23 hours of audio from 1,446 human-human conversations between 59 unique speakers. We…

Machine Learning · Computer Science 2021-03-22 Mike Wu , Jonathan Nafziger , Anthony Scodary , Andrew Maas

We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.6M) characters, and a vocabulary of 295,433 (295.4K) unique word types. We assembled the…

Computation and Language · Computer Science 2026-04-14 Haq Nawaz Malik , Nahfid Nissar

The following paper presents a project focused on the research and creation of a new Automatic Speech Recognition (ASR) based in the Chukchi language. There is no one complete corpus of the Chukchi language, so most of the work consisted in…

Computation and Language · Computer Science 2022-10-13 Anastasia Safonova , Tatiana Yudina , Emil Nadimanov , Cydnie Davenport

This paper introduces FT Speech, a new speech corpus created from the recorded meetings of the Danish Parliament, otherwise known as the Folketing (FT). The corpus contains over 1,800 hours of transcribed speech by a total of 434 speakers.…

Computation and Language · Computer Science 2020-10-29 Andreas Kirkedal , Marija Stepanović , Barbara Plank

Speech datasets available in the public domain are often underutilized because of challenges in discoverability and interoperability. A comprehensive framework has been designed to survey, catalog, and curate available speech datasets,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-02 Michał Junczyk

We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40…

Computation and Language · Computer Science 2026-04-02 Mohammad Mohammadamini , Daban Q. Jaff , Josep Crego , Marie Tahon , Antoine Laurent

Ramsa is a developing 41-hour speech corpus of Emirati Arabic designed to support sociolinguistic research and low-resource language technologies. It contains recordings from structured interviews with native speakers and episodes from…

Computation and Language · Computer Science 2026-03-10 Rania Al-Sabbagh

Existing Indic ASR benchmarks often use scripted, clean speech and leaderboard driven evaluation that encourages dataset specific overfitting. In addition, strict single reference WER penalizes natural spelling variation in Indian…

We present the development of a dataset for Kazakh named entity recognition. The dataset was built as there is a clear need for publicly available annotated corpora in Kazakh, as well as annotation guidelines containing straightforward--but…

Computation and Language · Computer Science 2022-04-08 Rustem Yeshpanov , Yerbolat Khassanov , Huseyin Atakan Varol

For conversational large-vocabulary continuous speech recognition (LVCSR) tasks, up to about two thousand hours of audio is commonly used to train state of the art models. Collection of labeled conversational audio however, is prohibitively…

Computation and Language · Computer Science 2017-05-30 Shane Walker , Morten Pedersen , Iroro Orife , Jason Flaks
‹ Prev 1 4 5 6 7 8 10 Next ›