English
Related papers

Related papers: A Crowdsourced Open-Source Kazakh Speech Corpus an…

200 papers

This article describes the MyST corpus developed as part of the My Science Tutor project -- one of the largest collections of children's conversational speech comprising approximately 400 hours, spanning some 230K utterances across about…

Computation and Language · Computer Science 2023-09-26 Sameer S. Pradhan , Ronald A. Cole , Wayne H. Ward

Automatic speech recognition (ASR) systems are designed to transcribe spoken language into written text and find utility in a variety of applications including voice assistants and transcription services. However, it has been observed that…

Computation and Language · Computer Science 2023-07-21 Anand Kumar Rai , Siddharth D Jaiswal , Animesh Mukherjee

This technical report presents the 600K-KS-OCR Dataset, a large-scale synthetic corpus comprising approximately 602,000 word-level segmented images designed for training and evaluating optical character recognition systems targeting…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Haq Nawaz Malik

This paper presents a high quality Vietnamese speech corpus that can be used for analyzing Vietnamese speech characteristic as well as building speech synthesis models. The corpus consists of 5400 clean-speech utterances spoken by 12…

Computation and Language · Computer Science 2019-04-12 Pham Ngoc Phuong , Quoc Truong Do , Luong Chi Mai

Thanks to improvements in machine learning techniques including deep learning, a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role. However, such a corpus for…

Computation and Language · Computer Science 2017-11-02 Ryosuke Sonobe , Shinnosuke Takamichi , Hiroshi Saruwatari

Speech recognition has received a less attention in Bengali literature due to the lack of a comprehensive dataset. In this paper, we describe the development process of the first comprehensive Bengali speech dataset on real numbers. It…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-28 Md Mahadi Hasan Nahid , Md. Ashraful Islam , Bishwajit Purkaystha , Md Saiful Islam

The lack of impaired speech data hinders advancements in the development of inclusive speech technologies, particularly in low-resource languages such as Akan. To address this gap, this study presents a curated corpus of speech samples from…

This paper introduces the development of the first open conversational speech dataset for the Isan language, the most widely spoken regional dialect in Thailand. Unlike existing speech corpora that are primarily based on read or scripted…

Computation and Language · Computer Science 2025-12-05 Adisai Na-Thalang , Chanakan Wittayasakpan , Kritsadha Phatcharoen , Supakit Buakaw

Most speech and language technologies are trained with massive amounts of speech and text information. However, most of the world languages do not have such resources or stable orthography. Systems constructed under these almost zero…

This paper presents an augmentation of MSCOCO dataset where speech is added to image and text. Speech captions are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.…

Computation and Language · Computer Science 2020-11-24 William Havard , Laurent Besacier , Olivier Rosec

We present STT4SG-350 (Speech-to-Text for Swiss German), a corpus of Swiss German speech, annotated with Standard German text at the sentence level. The data is collected using a web app in which the speakers are shown Standard German…

We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to talk shows and…

Computation and Language · Computer Science 2025-09-25 Samuel Stucki , Mark Cieliebak , Jan Deriu

Spontaneous conversations in real-world settings such as those found in child-centered recordings have been shown to be amongst the most challenging audio files to process. Nevertheless, building speech processing models handling such a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Marvin Lavechin , Ruben Bousbib , Hervé Bredin , Emmanuel Dupoux , Alejandrina Cristia

The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR resources for Puno Quechua (ISO 639-3: qxp): (1) the largest speech corpus for any…

Computation and Language · Computer Science 2026-05-28 Elwin Huaman , Adrian Gamarra Lafuente , Johanna Cordova , Anna Korhonen

A balanced speech corpus is the basic need for any speech processing task. In this report we describe our effort on development of Assamese speech corpus. We mainly focused on some issues and challenges faced during development of the…

Computation and Language · Computer Science 2013-09-30 Himangshu Sarma , Navanath Saharia , Utpal Sharma , Smriti Kumar Sinha , Mancha Jyoti Malakar

This paper introduces a new open-source speech corpus named "speechocean762" designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are children. Five…

Computation and Language · Computer Science 2021-06-03 Junbo Zhang , Zhiwen Zhang , Yongqing Wang , Zhiyong Yan , Qiong Song , Yukai Huang , Ke Li , Daniel Povey , Yujun Wang

The increase in technological adoption worldwide comes with demands for novel tools to be used by the general population. Large Language Models (LLMs) provide a great opportunity in this respect, but their capabilities remain limited for…

Computation and Language · Computer Science 2025-10-13 Stefan Krsteski , Matea Tashkovska , Borjan Sazdov , Hristijan Gjoreski , Branislav Gerazov

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

Code-switching refers to the usage of two languages within a sentence or discourse. It is a global phenomenon among multilingual communities and has emerged as an independent area of research. With the increasing demand for the…

Computation and Language · Computer Science 2018-10-02 Ganji Sreeram , Kunal Dhawan , Rohit Sinha

This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages. It contains two collections of datasets: unlabelled audio recordings of radio news and talk shows programs (160 hours) and labelled data…