中文
相关论文

相关论文: Developing a Fine-Grained Corpus for a Less-resour…

200 篇论文

Language recognition has been significantly advanced in recent years by means of modern machine learning methods such as deep learning and benchmarks with rich annotations. However, research is still limited in low-resource formal…

计算与语言 · 计算机科学 2020-06-03 Hadi Abdi Khojasteh , Ebrahim Ansari , Mahdi Bohlouli

We present an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. Surveying language documentation corpora and other resources that cover 67 languages and varieties…

计算与语言 · 计算机科学 2022-05-11 Andreas Liesenfeld , Mark Dingemanse

Discourse parsing is an important task useful for NLU applications such as summarization, machine comprehension, and emotion recognition. The current discourse parsing datasets based on conversations consists of written English dialogues…

计算与语言 · 计算机科学 2025-06-11 Divyaksh Shukla , Ritesh Baviskar , Dwijesh Gohil , Aniket Tiwari , Atul Shree , Ashutosh Modi

Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs). While various efforts have been made to collect monolingual and multilingual datasets in many languages, Persian…

This paper presents a comprehensive survey of corpora and lexical resources available for Turkish. We review a broad range of resources, focusing on the ones that are publicly available. In addition to providing information about the…

计算与语言 · 计算机科学 2023-02-28 Çağrı Çöltekin , A. Seza Doğruöz , Özlem Çetinoğlu

Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. The absence of large-scale, high-quality datasets has limited the development of Urdu-capable systems and reinforced biases…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Umair Hassan

We introduce the Tarab Corpus, a large-scale cultural and linguistic resource that brings together Arabic song lyrics and poetry within a unified analytical framework. The corpus comprises 2.56 million verses and more than 13.5 million…

计算与语言 · 计算机科学 2026-03-18 Mo El-Haj

The need for large text corpora has increased with the advent of pretrained language models and, in particular, the discovery of scaling laws for these models. Most available corpora have sufficient data only for languages with large…

计算与语言 · 计算机科学 2025-03-05 Amir Hossein Kargaran , François Yvon , Hinrich Schütze

We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. The corpora are…

计算与语言 · 计算机科学 2020-07-16 Shashank Siripragada , Jerin Philip , Vinay P. Namboodiri , C V Jawahar

Lectures translation is a case of spoken language translation and there is a lack of publicly available parallel corpora for this purpose. To address this, we examine a language independent framework for parallel corpus mining which is a…

计算与语言 · 计算机科学 2020-01-15 Haiyue Song , Raj Dabre , Atsushi Fujita , Sadao Kurohashi

At present, Text-to-speech (TTS) systems that are trained with high-quality transcribed speech data using end-to-end neural models can generate speech that is intelligible, natural, and closely resembles human speech. These models are…

计算与语言 · 计算机科学 2023-03-02 Ajinkya Kulkarni , Atharva Kulkarni , Sara Abedalmonem Mohammad Shatnawi , Hanan Aldarmaki

Current large language models demonstrate deficiencies in understanding low-resource languages, particularly the minority languages in China. This limitation stems from the scarcity of available pre-training data. To address this…

计算与语言 · 计算机科学 2024-06-14 Chen Zhang , Mingxu Tao , Quzhe Huang , Jiuheng Lin , Zhibin Chen , Yansong Feng

Research in NLP for Central Asian Turkic languages - Kazakh, Uzbek, Kyrgyz, and Turkmen - faces typical low-resource language challenges like data scarcity, limited linguistic resources and technology development. However, recent…

计算与语言 · 计算机科学 2026-02-17 Yana Veitsman , Mareike Hartmann

Over the past years, interest in discourse analysis and discourse parsing has steadily grown, and many discourse-annotated corpora and, as a result, discourse parsers have been built. In this paper, we present a discourse-annotated corpus…

计算与语言 · 计算机科学 2021-06-29 Sara Shahmohammadi , Hadi Veisi , Ali Darzi

End-to-end transformer-based models epitomize the cutting-edge in Automatic Speech Recognition (ASR) systems. Despite their substantial benefits, these models demand extensive training data to perform optimally, presenting a significant…

音频与语音处理 · 电气工程与系统科学 2024-09-10 Abdulhady Abas Abdullah , Hadi Veisi , Tarik Rashid

In this paper, we introduce the French-YMCA corpus, a new linguistic resource specifically tailored for children and adolescents. The motivation for building this corpus is clear: children have unique language requirements, as their…

计算与语言 · 计算机科学 2026-04-08 Cherifa Ben Khelil , Jean-Yves Antoine , Anaïs Halftermeyer , Frédéric Rayar , Mathieu Thebaud

Sanskrit is a classical language with about 30 million extant manuscripts fit for digitisation, available in written, printed or scannedimage forms. However, it is still considered to be a low-resource language when it comes to available…

计算与语言 · 计算机科学 2022-11-16 Ayush Maheshwari , Nikhil Singh , Amrith Krishna , Ganesh Ramakrishnan

Despite remarkable progress in large language models, Urdu-a language spoken by over 230 million people-remains critically underrepresented in modern NLP systems. Existing multilingual models demonstrate poor performance on Urdu-specific…

计算与语言 · 计算机科学 2026-01-14 Muhammad Taimoor Hassan , Jawad Ahmed , Muhammad Awais

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article…

This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two…

音频与语音处理 · 电气工程与系统科学 2021-09-09 Saida Mussakhojayeva , Aigerim Janaliyeva , Almas Mirzakhmetov , Yerbolat Khassanov , Huseyin Atakan Varol