中文
相关论文

相关论文: Building Community-Centred NLP Resources for Puno …

200 篇论文

Spoken Language Understanding (SLU) aims to extract the semantic information from the speech utterance of user queries. It is a core component in a task-oriented dialogue system. With the spectacular progress of deep neural network models…

计算与语言 · 计算机科学 2026-03-24 Haroun Elleuch , Salima Mdhaffar , Yannick Estève , Fethi Bougares

We present LOTUSDIS, a publicly available Thai meeting corpus designed to advance far-field conversational ASR. The dataset comprises 114 hours of spontaneous, unscripted dialogue collected in 15-20 minute sessions with three participants,…

计算与语言 · 计算机科学 2025-09-24 Pattara Tipaksorn , Sumonmas Thatphithakkul , Vataya Chunwijitra , Kwanchiva Thangthai

With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus…

音频与语音处理 · 电气工程与系统科学 2024-06-21 Linhan Ma , Dake Guo , Kun Song , Yuepeng Jiang , Shuai Wang , Liumeng Xue , Weiming Xu , Huan Zhao , Binbin Zhang , Lei Xie

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children's speech…

While the NLP community is generally aware of resource disparities among languages, we lack research that quantifies the extent and types of such disparity. Prior surveys estimating the availability of resources based on the number of…

计算与语言 · 计算机科学 2022-11-29 Xinyan Velocity Yu , Akari Asai , Trina Chatterjee , Junjie Hu , Eunsol Choi

Self-supervised learning (SSL) has helped extend speech technologies to more languages by reducing the need for labeled data. However, models are still far from supporting the world's 7000+ languages. We propose XEUS, a Cross-lingual…

The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue is exacerbated in the era of large, data-intensive models.…

计算与语言 · 计算机科学 2024-06-24 Georgios Paraskevopoulos , Chara Tsoukala , Athanasios Katsamanis , Vassilis Katsouros

For many of the 700 million illiterate people around the world, speech recognition technology could provide a bridge to valuable information and services. Yet, those most in need of this technology are often the most underserved by it. In…

机器学习 · 计算机科学 2021-04-28 Moussa Doumbouya , Lisa Einstein , Chris Piech

This technical report describes the results of a collaboration between the NLP research group at the University of Tartu and the Institute of Estonian Language on improving neural speech synthesis for Estonian. The report (written in…

计算与语言 · 计算机科学 2020-10-07 Liisa Rätsep , Liisi Piits , Hille Pajupuu , Indrek Hein , Mark Fišel

This paper introduces Multilingual LibriSpeech (MLS) dataset, a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages, including about 44.5K hours of…

音频与语音处理 · 电气工程与系统科学 2020-12-22 Vineel Pratap , Qiantong Xu , Anuroop Sriram , Gabriel Synnaeve , Ronan Collobert

Automatic speech recognition systems have undoubtedly advanced with the integration of multilingual and multitask models such as Whisper, which have shown a promising ability to understand and process speech across a wide range of…

计算与语言 · 计算机科学 2025-04-14 Xabier de Zuazo , Eva Navas , Ibon Saratxaga , Inma Hernáez Rioja

Slovak remains a low-resource language for automatic speech recognition (ASR), with fewer than 100 hours of publicly available training data. We present SloPal, a comprehensive Slovak parliamentary corpus comprising 330,000…

计算与语言 · 计算机科学 2026-03-17 Erik Božík , Marek Šuppa

St. Lawrence Island Yupik (ISO 639-3: ess) is an endangered polysynthetic language in the Inuit-Yupik language family indigenous to Alaska and Chukotka. This work presents a step-by-step pipeline for the digitization of written texts, and…

计算与语言 · 计算机科学 2021-01-27 Lane Schwartz , Emily Chen , Hyunji Hayley Park , Edward Jahn , Sylvia L. R. Schreiner

We present the Multilingual Cloud Corpus, the first national-scale, parallel, multimodal linguistic dataset of Bangladesh's ethnic and indigenous languages. Despite being home to approximately 40 minority languages spanning four language…

计算与语言 · 计算机科学 2026-03-09 Mohammad Mamun Or Rashid

Recent advances in Automatic Speech Recognition (ASR) have been largely fueled by massive speech corpora. However, extending coverage to diverse languages with limited resources remains a formidable challenge. This paper introduces Speech…

计算与语言 · 计算机科学 2025-05-23 Tianduo Wang , Lu Xu , Wei Lu , Shanbo Cheng

Automatic Speech Recognition (ASR) has increasing utility in the modern world. There are a many ASR models available for languages with large amounts of training data like English. However, low-resource languages are poorly represented. In…

音频与语音处理 · 电气工程与系统科学 2023-05-24 Kavitha Raju , Anjaly V , Ryan Lish , Joel Mathew

Spoken query retrieval is an important interaction mode in modern information retrieval. However, existing evaluation datasets are often limited to simple queries under constrained noise conditions, making them inadequate for assessing the…

信息检索 · 计算机科学 2026-05-14 Yuejie Li , Ke Yang , Yueying Hua , Berlin Chen , Jianhao Nie , Yueping He , Caixin Kang

Despite 230 million speakers, Urdu remains critically under-resourced in speech technology. We introduce UrduSpeech: a large high-fidelity Urdu corpus comprising 156 hours of audio with 12-dimension paralinguistic metadata, encompassing…

音频与语音处理 · 电气工程与系统科学 2026-05-19 Attia Nafees ul Haq , Zeyu Zhu , Jingbin Hu , ChunJiang He , Lei Xie

Speech technology remains out of reach for most of the over 2300 languages in Africa. We present the first systematic assessment of large-scale synthetic voice corpora for African ASR. We apply a three-step process: LLM-driven text…

计算与语言 · 计算机科学 2025-11-10 Brian DeRenzi , Anna Dixon , Mohamed Aymane Farhi , Christian Resch

We present HebDB, a weakly supervised dataset for spoken language processing in the Hebrew language. HebDB offers roughly 2500 hours of natural and spontaneous speech recordings in the Hebrew language, consisting of a large variety of…