English
Related papers

Related papers: Building Community-Centred NLP Resources for Puno …

200 papers

Despite significant advances in speech processing, Portuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over…

There is a major shortage of Speech-to-Speech Translation (S2ST) datasets for high resource-to-low resource language pairs such as English-to-Yoruba. Thus, in this study, we curated the Bilingual English-to-Yoruba Speech-to-Speech…

We present Ara-BEST-RQ, a family of self-supervised learning (SSL) models specifically designed for multi-dialectal Arabic speech processing. Leveraging 5,640 hours of crawled Creative Commons speech and combining it with publicly available…

Computation and Language · Computer Science 2026-03-24 Haroun Elleuch , Ryan Whetten , Salima Mdhaffar , Yannick Estève , Fethi Bougares

Natural language processing (NLP) and speech technologies have made significant progress in recent years; however, they remain largely focused on standardized language varieties. Dialects, despite their cultural significance and widespread…

Computation and Language · Computer Science 2026-04-14 Lena S. Oberkircher , Jesujoba O. Alabi , Dietrich Klakow , Jürgen Trouvain

The digital exclusion of endangered languages remains a critical challenge in NLP, limiting both linguistic research and revitalization efforts. This study introduces the first computational investigation of Comanche, an Uto-Aztecan…

Computation and Language · Computer Science 2025-05-27 Jesus Alvarez C , Daua D. Karajeanes , Ashley Celeste Prado , John Ruttan , Ivory Yang , Sean O'Brien , Vasu Sharma , Kevin Zhu

Large language models (LLMs) have driven substantial advances in speech language models (SpeechLMs), yielding strong performance in automatic speech recognition (ASR) under high-resource conditions. However, existing benchmarks…

Computation and Language · Computer Science 2026-03-23 Jianan Chen , Xiaoxue Gao , Tatsuya Kawahara , Nancy F. Chen

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted…

Automatic speech recognition (ASR) for African languages remains constrained by limited labeled data and the lack of systematic guidance on model selection, data scaling, and decoding strategies. Large pre-trained systems such as Whisper,…

The recent advances in natural language processing (NLP) are linked to training processes that require vast amounts of corpora. Access to this data is commonly not a trivial process due to resource dispersion and the need to maintain these…

Computation and Language · Computer Science 2024-01-30 Rúben Almeida , Ricardo Campos , Alípio Jorge , Sérgio Nunes

Speech large language models (SLLMs) built on speech encoders, adapters, and LLMs demonstrate remarkable multitask understanding performance in high-resource languages such as English and Chinese. However, their effectiveness substantially…

Sound · Computer Science 2026-04-21 Mingchen Shao , Bingshen Mu , Chengyou Wang , Hai Li , Ying Yan , Zhonghua Fu , Lei Xie

This preprint describes work in progress on LR-Sum, a new permissively-licensed dataset created with the goal of enabling further research in automatic summarization for less-resourced languages. LR-Sum contains human-written summaries for…

Computation and Language · Computer Science 2023-10-30 Chester Palen-Michel , Constantine Lignos

Creating speech datasets for low-resource languages is a critical yet poorly understood challenge, particularly regarding the actual cost in human labor. This paper investigates the time and complexity required to produce high-quality…

Computation and Language · Computer Science 2025-10-15 Yacouba Diarra , Nouhoum Souleymane Coulibaly , Michael Leventhal

Democratizing access to natural language processing (NLP) technology is crucial, especially for underrepresented and extremely low-resource languages. Previous research has focused on developing labeled and unlabeled corpora for these…

We present a freely available speech corpus for the Uzbek language and report preliminary automatic speech recognition (ASR) results using both the deep neural network hidden Markov model (DNN-HMM) and end-to-end (E2E) architectures. The…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-02 Muhammadjon Musaev , Saida Mussakhojayeva , Ilyos Khujayorov , Yerbolat Khassanov , Mannon Ochilov , Huseyin Atakan Varol

Whisper generation is constrained by the difficulty of data collection. Because whispered speech has low acoustic amplitude, high-fidelity recording is challenging. In this paper, we introduce WhispSynth, a large-scale multilingual corpus…

Sound · Computer Science 2026-03-17 Tianyi Tan , Jiaxin Ye , Yuanming Zhang , Xiaohuai Le , Xianjun Xia , Chuanzeng Huang , Jing Lu

An open-source Mandarin speech corpus called AISHELL-1 is released. It is by far the largest corpus which is suitable for conducting the speech recognition research and building speech recognition systems for Mandarin. The recording…

Computation and Language · Computer Science 2017-09-19 Hui Bu , Jiayu Du , Xingyu Na , Bengu Wu , Hao Zheng

Building automatic speech recognition (ASR) systems is a challenging task, especially for under-resourced languages that need to construct corpora nearly from scratch and lack sufficient training data. It has emerged that several African…

Computation and Language · Computer Science 2022-11-01 Ebbie Awino , Lilian Wanzare , Lawrence Muchemi , Barack Wanjawa , Edward Ombui , Florence Indede , Owen McOnyango , Benard Okal

Listening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings into long-form spoken language documents, reading ASR…

Computation and Language · Computer Science 2023-03-28 Qinglin Zhang , Chong Deng , Jiaqing Liu , Hai Yu , Qian Chen , Wen Wang , Zhijie Yan , Jinglin Liu , Yi Ren , Zhou Zhao

This paper introduces a centralized, open-source dataset repository designed to advance NLP and NMT for Assamese, a low-resource language. The repository, available at GitHub, supports various tasks like sentiment analysis, named entity…

Computation and Language · Computer Science 2024-10-17 S. Tamang , D. J. Bora