English
Related papers

Related papers: FeruzaSpeech: A 60 Hour Uzbek Read Speech Corpus w…

200 papers

Despite Telugu being spoken by over 80 million people, speech translation research for this morphologically rich language remains severely underexplored. We address this gap by developing a high-quality Telugu--English speech translation…

Computation and Language · Computer Science 2025-12-09 Bhavana Akkiraju , Srihari Bandarupalli , Swathi Sambangi , Vasavi Ravuri , R Vijaya Saraswathi , Anil Kumar Vuppala

Keyphrases provide an extremely dense summary of a text. Such information can be used in many Natural Language Processing tasks, such as information retrieval and text summarization. Since previous studies on Persian keyword or keyphrase…

Computation and Language · Computer Science 2020-09-28 Ehsan Doostmohammadi , Mohammad Hadi Bokaei , Hossein Sameti

In this work, we showcase a cost-effective method for generating training data for speech processing tasks. First, we transcribe unlabeled speech using a state-of-the-art Automatic Speech Recognition (ASR) model. Next, we align generated…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-19 Taras Sereda

Extracting useful information for sentiment analysis and classification problems from a big amount of user-generated feedback, such as restaurant reviews, is a crucial task of natural language processing, which is not only for customer…

Computation and Language · Computer Science 2022-06-01 Sanatbek Matlatipov , Hulkar Rahimboeva , Jaloliddin Rajabov , Elmurod Kuriyozov

In this paper, we introduce the first large vocabulary speech recognition system (LVSR) for the Central Kurdish language, named Jira. The Kurdish language is an Indo-European language spoken by more than 30 million people in several…

Artificial Intelligence · Computer Science 2021-02-16 Hadi Veisi , Hawre Hosseini , Mohammad Mohammadamini , Wirya Fathy , Aso Mahmudi

Speech recognition has received a less attention in Bengali literature due to the lack of a comprehensive dataset. In this paper, we describe the development process of the first comprehensive Bengali speech dataset on real numbers. It…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-28 Md Mahadi Hasan Nahid , Md. Ashraful Islam , Bishwajit Purkaystha , Md Saiful Islam

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one…

Computation and Language · Computer Science 2022-06-28 Chengfei Li , Shuhao Deng , Yaoping Wang , Guangjing Wang , Yaguang Gong , Changbin Chen , Jinfeng Bai

DeepMine is a speech database in Persian and English designed to build and evaluate text-dependent, text-prompted, and text-independent speaker verification, as well as Persian speech recognition systems. It contains more than 1850 speakers…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-10 Hossein Zeinali , Lukáš Burget , Jan "Honza'' Černocký

This paper introduces a set of English translations for a 123-hour subset of the CallHome Mandarin Chinese data and the HKUST Mandarin Telephone Speech data for the task of speech translation. Paired source-language speech and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-19 Shannon Wotherspoon , William Hartmann , Matthew Snover

In this paper we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. These two datasets were made available for the IWSLT 2022 low-resource speech translation track, and they consist of collections of…

Computation and Language · Computer Science 2022-04-12 Marcely Zanon Boito , Fethi Bougares , Florentin Barbier , Souhir Gahbiche , Loïc Barrault , Mickael Rouvier , Yannick Estève

We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40…

Computation and Language · Computer Science 2026-04-02 Mohammad Mohammadamini , Daban Q. Jaff , Josep Crego , Marie Tahon , Antoine Laurent

This study focuses on the creation of the KazEmoTTS dataset, designed for emotional Kazakh text-to-speech (TTS) applications. KazEmoTTS is a collection of 54,760 audio-text pairs, with a total duration of 74.85 hours, featuring 34.23 hours…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-11 Adal Abilbekov , Saida Mussakhojayeva , Rustem Yeshpanov , Huseyin Atakan Varol

This paper presents a novel Dialectal Sound and Vowelization Recovery framework, designed to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages, that extends beyond its standard orthographic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-06 Yassine El Kheir , Hamdy Mubarak , Ahmed Ali , Shammur Absar Chowdhury

Documenting languages helps to prevent the extinction of endangered dialects, many of which are otherwise expected to disappear by the end of the century. When documenting oral languages, unsupervised word segmentation (UWS) from speech is…

Computation and Language · Computer Science 2022-05-19 Marcely Zanon Boito , Bolaji Yusuf , Lucas Ondel , Aline Villavicencio , Laurent Besacier

In this paper, we construct a new Japanese speech corpus called "JTubeSpeech." Although recent end-to-end learning requires large-size speech corpora, open-sourced such corpora for languages other than English have not yet been established.…

Despite growing interest in Quranic data research, existing Quran datasets remain limited in both scale and diversity. To address this gap, we present Tadabur, a large-scale Quran audio dataset. Tadabur comprises more than 1400+ hours of…

Sound · Computer Science 2026-04-22 Faisal Alherran

This paper presents an extension to a very low-resource parallel corpus collected in an endangered language, Griko, making it useful for computational research. The corpus consists of 330 utterances (about 20 minutes of speech) which have…

Computation and Language · Computer Science 2018-07-30 Marcely Zanon Boito , Antonios Anastasopoulos , Marika Lekakou , Aline Villavicencio , Laurent Besacier

We present a test corpus of audio recordings and transcriptions of presentations of students' enterprises together with their slides and web-pages. The corpus is intended for evaluation of automatic speech recognition (ASR) systems,…

Computation and Language · Computer Science 2019-08-05 Dominik Macháček , Jonáš Kratochvíl , Tereza Vojtěchová , Ondřej Bojar

This paper introduces Mixat: a dataset of Emirati speech code-mixed with English. Mixat was developed to address the shortcomings of current speech recognition resources when applied to Emirati speech, and in particular, to bilignual…

Computation and Language · Computer Science 2024-05-07 Maryam Al Ali , Hanan Aldarmaki

This study addresses automatic transliteration from Tajik (Cyrillic script) to Persian (Perso-Arabic script). We present a curated, lexicographically verified parallel corpus of 52,152 Tajik--Persian words and short phrases, compiled from…

Computation and Language · Computer Science 2026-05-12 Mullosharaf K. Arabov