English
Related papers

Related papers: Language and Speech Technology for Central Kurdish…

200 papers

Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. The absence of large-scale, high-quality datasets has limited the development of Urdu-capable systems and reinforced biases…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Umair Hassan

We present an experimental dataset, Basic Dataset for Sorani Kurdish Automatic Speech Recognition (BD-4SK-ASR), which we used in the first attempt in developing an automatic speech recognition for Sorani Kurdish. The objective of the…

Computation and Language · Computer Science 2019-12-03 Akam Qader , Hossein Hassani

We present an open-source speech corpus for the Kazakh language. The Kazakh speech corpus (KSC) contains around 332 hours of transcribed audio comprising over 153,000 utterances spoken by participants from different regions and age groups,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-22 Yerbolat Khassanov , Saida Mussakhojayeva , Almas Mirzakhmetov , Alen Adiyev , Mukhamet Nurpeiissov , Huseyin Atakan Varol

- The field of natural language processing (NLP) has dramatically expanded within the last decade. Many human-being applications are conducted daily via NLP tasks, starting from machine translation, speech recognition, text generation and…

Classifying Sorani Kurdish subdialects poses a challenge due to the need for publicly available datasets or reliable resources like social media or websites for data collection. We conducted field visits to various cities and villages to…

Computation and Language · Computer Science 2024-04-02 Sana Isam , Hossein Hassani

Research in NLP for Central Asian Turkic languages - Kazakh, Uzbek, Kyrgyz, and Turkmen - faces typical low-resource language challenges like data scarcity, limited linguistic resources and technology development. However, recent…

Computation and Language · Computer Science 2026-02-17 Yana Veitsman , Mareike Hartmann

Yor\`ub\'a an African language with roughly 47 million speakers encompasses a continuum with several dialects. Recent efforts to develop NLP technologies for African languages have focused on their standard dialects, resulting in…

Computation and Language · Computer Science 2024-07-01 Orevaoghene Ahia , Anuoluwapo Aremu , Diana Abagyan , Hila Gonen , David Ifeoluwa Adelani , Daud Abolade , Noah A. Smith , Yulia Tsvetkov

At present, Text-to-speech (TTS) systems that are trained with high-quality transcribed speech data using end-to-end neural models can generate speech that is intelligible, natural, and closely resembles human speech. These models are…

Computation and Language · Computer Science 2023-03-02 Ajinkya Kulkarni , Atharva Kulkarni , Sara Abedalmonem Mohammad Shatnawi , Hanan Aldarmaki

We develop machine translation and speech synthesis systems to complement the efforts of revitalizing Judeo-Spanish, the exiled language of Sephardic Jews, which survived for centuries, but now faces the threat of extinction in the digital…

Computation and Language · Computer Science 2022-06-01 Alp Öktem , Rodolfo Zevallos , Yasmin Moslem , Güneş Öztürk , Karen Şarhon

Hawrami, a dialect of Kurdish, is classified as an endangered language as it suffers from the scarcity of data and the gradual loss of its speakers. Natural Language Processing projects can be used to partially compensate for data…

Computation and Language · Computer Science 2024-09-26 Aram Khaksar , Hossein Hassani

Automatic Speech Recognition (ASR) for low-resource languages remains a challenging task due to limited training data. This paper introduces a comprehensive study exploring the effectiveness of Whisper, a pre-trained ASR model, for Northern…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-23 Abdulhady Abas Abdullah , Shima Tabibian , Hadi Veisi , Aso Mahmudi , Tarik Rashid

We present an expanded version of our previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In the new KazakhTTS2 corpus, the overall size has increased from 93 hours to 271 hours, the number of speakers has risen from two…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-21 Saida Mussakhojayeva , Yerbolat Khassanov , Huseyin Atakan Varol

Bridging linguistic gaps fosters global growth and cultural exchange. This study addresses the challenges of Roman Urdu -- a Latin-script adaptation of Urdu widely used in digital communication -- by creating a novel parallel dataset…

Computation and Language · Computer Science 2024-12-24 Mohammed Furqan , Raahid Bin Khaja , Rayyan Habeeb

In this study, we develop and assess new corpus selection and training methodologies to improve the effectiveness of Turkish language models. Specifically, we adapted Large Language Model generated datasets and translated English datasets…

Computation and Language · Computer Science 2024-12-05 H. Toprak Kesgin , M. Kaan Yuce , Eren Dogan , M. Egemen Uzun , Atahan Uz , Elif Ince , Yusuf Erdem , Osama Shbib , Ahmed Zeer , M. Fatih Amasyali

FLEURS offers n-way parallel speech for 100+ languages, but Northern Kurdish is not one of them, which limits benchmarking for automatic speech recognition and speech translation tasks in this language. We present FLEURS-Kobani, a Northern…

Computation and Language · Computer Science 2026-04-01 Daban Q. Jaff , Mohammad Mohammadamini

Access to Kurdish medicine brochures is limited, depriving Kurdish-speaking communities of critical health information. To address this problem, we developed a specialized Machine Translation (MT) model to translate English medicine…

Computation and Language · Computer Science 2025-01-24 Mariam Shamal , Hossein Hassani

Idiom detection using Natural Language Processing (NLP) is the computerized process of recognizing figurative expressions within a text that convey meanings beyond the literal interpretation of the words. While idiom detection has seen…

Computation and Language · Computer Science 2025-08-19 Skala Kamaran Omer , Hossein Hassani

Dubbed series are gaining a lot of popularity in recent years with strong support from major media service providers. Such popularity is fueled by studies that showed that dubbed versions of TV shows are more popular than their subtitled…

Computation and Language · Computer Science 2022-03-08 Massa Baali , Wassim El-Hajj , Ahmed Ali

This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-09 Saida Mussakhojayeva , Aigerim Janaliyeva , Almas Mirzakhmetov , Yerbolat Khassanov , Huseyin Atakan Varol

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source…

Computation and Language · Computer Science 2021-06-16 Elizabeth Salesky , Matthew Wiesner , Jacob Bremerman , Roldano Cattoni , Matteo Negri , Marco Turchi , Douglas W. Oard , Matt Post