中文
相关论文

相关论文: A Crowdsourced Open-Source Kazakh Speech Corpus an…

200 篇论文

For many of the 700 million illiterate people around the world, speech recognition technology could provide a bridge to valuable information and services. Yet, those most in need of this technology are often the most underserved by it. In…

机器学习 · 计算机科学 2021-04-28 Moussa Doumbouya , Lisa Einstein , Chris Piech

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce…

Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt dataset for safety evaluation across eleven categories covering common risk areas such as…

计算与语言 · 计算机科学 2026-05-29 Wajdi Zaghouani , Shimaa Amer Ibrahim , Aruzhan Muratbek , Olzhasbek Zhakenov , Adiya Akhmetzhanova

Whisper and other large-scale automatic speech recognition models have made significant progress in performance. However, their performance on many low-resource languages, such as Kazakh, is not satisfactory. It is worth researching how to…

音频与语音处理 · 电气工程与系统科学 2025-05-08 Jinpeng Li , Yu Pu , Qi Sun , Wei-Qiang Zhang

Parliamentary transcripts provide a valuable resource to understand the reality and know about the most important facts that occur over time in our societies. Furthermore, the political debates captured in these transcripts facilitate…

Recent advances in audio-language models have demonstrated remarkable success on short, segment-level speech tasks. However, real-world applications such as meeting transcription, spoken document understanding, and conversational analysis…

We introduce VoxPopuli, a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages. It is the largest open data to date for unsupervised representation learning as well as semi-supervised learning.…

In this paper, we present a new Russian and Kazakh database (with about 95% of Russian and 5% of Kazakh words/sentences respectively) for offline handwriting recognition. A few pre-processing and segmentation procedures have been developed…

计算机视觉与模式识别 · 计算机科学 2021-08-31 Daniyar Nurseitov , Kairat Bostanbekov , Daniyar Kurmankhojayev , Anel Alimova , Abdelrahman Abdallah

Dubbed series are gaining a lot of popularity in recent years with strong support from major media service providers. Such popularity is fueled by studies that showed that dubbed versions of TV shows are more popular than their subtitled…

计算与语言 · 计算机科学 2022-03-08 Massa Baali , Wassim El-Hajj , Ahmed Ali

Slovak remains a low-resource language for automatic speech recognition (ASR), with fewer than 100 hours of publicly available training data. We present SloPal, a comprehensive Slovak parliamentary corpus comprising 330,000…

计算与语言 · 计算机科学 2026-03-17 Erik Božík , Marek Šuppa

We introduce HarperValleyBank, a free, public domain spoken dialog corpus. The data simulate simple consumer banking interactions, containing about 23 hours of audio from 1,446 human-human conversations between 59 unique speakers. We…

机器学习 · 计算机科学 2021-03-22 Mike Wu , Jonathan Nafziger , Anthony Scodary , Andrew Maas

We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.6M) characters, and a vocabulary of 295,433 (295.4K) unique word types. We assembled the…

计算与语言 · 计算机科学 2026-04-14 Haq Nawaz Malik , Nahfid Nissar

The following paper presents a project focused on the research and creation of a new Automatic Speech Recognition (ASR) based in the Chukchi language. There is no one complete corpus of the Chukchi language, so most of the work consisted in…

计算与语言 · 计算机科学 2022-10-13 Anastasia Safonova , Tatiana Yudina , Emil Nadimanov , Cydnie Davenport

This paper introduces FT Speech, a new speech corpus created from the recorded meetings of the Danish Parliament, otherwise known as the Folketing (FT). The corpus contains over 1,800 hours of transcribed speech by a total of 434 speakers.…

计算与语言 · 计算机科学 2020-10-29 Andreas Kirkedal , Marija Stepanović , Barbara Plank

Speech datasets available in the public domain are often underutilized because of challenges in discoverability and interoperability. A comprehensive framework has been designed to survey, catalog, and curate available speech datasets,…

音频与语音处理 · 电气工程与系统科学 2024-08-02 Michał Junczyk

We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40…

计算与语言 · 计算机科学 2026-04-02 Mohammad Mohammadamini , Daban Q. Jaff , Josep Crego , Marie Tahon , Antoine Laurent

Ramsa is a developing 41-hour speech corpus of Emirati Arabic designed to support sociolinguistic research and low-resource language technologies. It contains recordings from structured interviews with native speakers and episodes from…

计算与语言 · 计算机科学 2026-03-10 Rania Al-Sabbagh

Existing Indic ASR benchmarks often use scripted, clean speech and leaderboard driven evaluation that encourages dataset specific overfitting. In addition, strict single reference WER penalizes natural spelling variation in Indian…

We present the development of a dataset for Kazakh named entity recognition. The dataset was built as there is a clear need for publicly available annotated corpora in Kazakh, as well as annotation guidelines containing straightforward--but…

计算与语言 · 计算机科学 2022-04-08 Rustem Yeshpanov , Yerbolat Khassanov , Huseyin Atakan Varol

For conversational large-vocabulary continuous speech recognition (LVCSR) tasks, up to about two thousand hours of audio is commonly used to train state of the art models. Collection of labeled conversational audio however, is prohibitively…

计算与语言 · 计算机科学 2017-05-30 Shane Walker , Morten Pedersen , Iroro Orife , Jason Flaks