中文
相关论文

相关论文: A Study on the Appropriate size of the Mongolian g…

200 篇论文

This paper introduces a high-quality open-source text-to-speech (TTS) synthesis dataset for Mongolian, a low-resource language spoken by over 10 million people worldwide. The dataset, named MnTTS, consists of about 8 hours of transcribed…

声音 · 计算机科学 2022-09-23 Yifan Hu , Pengkai Yin , Rui Liu , Feilong Bao , Guanglai Gao

The increase in technological adoption worldwide comes with demands for novel tools to be used by the general population. Large Language Models (LLMs) provide a great opportunity in this respect, but their capabilities remain limited for…

计算与语言 · 计算机科学 2025-10-13 Stefan Krsteski , Matea Tashkovska , Borjan Sazdov , Hristijan Gjoreski , Branislav Gerazov

One of the major challenges of an educational system is choosing appropriate content considering pupils' age and intellectual potential. In this article the experiment of primary school grades (from 1st to 4th grades) is considered for…

计算与语言 · 计算机科学 2023-03-21 Khabibulla Madatov , Sanatbek Matlatipov , Mersaid Aripov

We present an expanded version of our previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In the new KazakhTTS2 corpus, the overall size has increased from 93 hours to 271 hours, the number of speakers has risen from two…

音频与语音处理 · 电气工程与系统科学 2022-04-21 Saida Mussakhojayeva , Yerbolat Khassanov , Huseyin Atakan Varol

Kurdish is a less-resourced language consisting of different dialects written in various scripts. Approximately 30 million people in different countries speak the language. The lack of corpora is one of the main obstacles in Kurdish…

计算与语言 · 计算机科学 2019-09-26 Roshna Omer Abdulrahman , Hossein Hassani , Sina Ahmadi

In this work, we introduce the MOldavian and ROmanian Dialectal COrpus (MOROCO), which is freely available for download at https://github.com/butnaruandrei/MOROCO. The corpus contains 33564 samples of text (with over 10 million tokens)…

计算与语言 · 计算机科学 2019-06-04 Andrei M. Butnaru , Radu Tudor Ionescu

Text-to-Speech (TTS) synthesis for low-resource languages is an attractive research issue in academia and industry nowadays. Mongolian is the official language of the Inner Mongolia Autonomous Region and a representative low-resource…

音频与语音处理 · 电气工程与系统科学 2023-01-03 Kailin Liang , Bin Liu , Yifan Hu , Rui Liu , Feilong Bao , Guanglai Gao

This paper introduces the hmBlogs corpus for Persian, as a low resource language. This corpus has been prepared based on a collection of nearly 20 million blog posts over a period of about 15 years from a space of Persian blogs and includes…

计算与语言 · 计算机科学 2021-11-04 Hamzeh Motahari Khansari , Mehrnoush Shamsfard

We present the Knesset Corpus, a corpus of Hebrew parliamentary proceedings containing over 30 million sentences (over 384 million tokens) from all the (plenary and committee) protocols held in the Israeli parliament between 1998 and 2022.…

计算与语言 · 计算机科学 2025-06-02 Gili Goldin , Nick Howell , Noam Ordan , Ella Rabinovich , Shuly Wintner

With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus…

音频与语音处理 · 电气工程与系统科学 2024-06-21 Linhan Ma , Dake Guo , Kun Song , Yuepeng Jiang , Shuai Wang , Liumeng Xue , Weiming Xu , Huan Zhao , Binbin Zhang , Lei Xie

A large amount of feedback was collected over the years. Many feedback analysis models have been developed focusing on the English language. Recognizing the concept of feedback is challenging and crucial in languages which do not have…

计算与语言 · 计算机科学 2023-02-24 Zolzaya Dashdorj , Tsetsentsengel Munkhbayar , Stanislav Grigorev

Mental health risk prediction is a growing field in the speech community, but many studies are based on small corpora. This study illustrates how variations in test and train set sizes impact performance in a controlled study. Using a…

计算与语言 · 计算机科学 2025-01-03 Tomek Rutowski , Amir Harati , Elizabeth Shriberg , Yang Lu , Piotr Chlebek , Ricardo Oliveira

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one…

计算与语言 · 计算机科学 2022-06-28 Chengfei Li , Shuhao Deng , Yaoping Wang , Guangjing Wang , Yaguang Gong , Changbin Chen , Jinfeng Bai

In this paper, we present a comprehensive corpus-driven analysis of Bangla literary and newspaper texts to investigate their lexical diversity, structural complexity and readability. We undertook Vacaspati and IndicCorp, which are the most…

计算与语言 · 计算机科学 2026-01-13 Pramit Bhattacharyya , Arnab Bhattacharya

This paper describes a web-based corpus of global language use with a focus on how this corpus can be used for data-driven language mapping. First, the corpus provides a representation of where national varieties of major languages are used…

计算与语言 · 计算机科学 2020-04-03 Jonathan Dunn

This paper proposes a novel framework for digital curation of Web corpora in order to provide robust estimation of their parameters, such as their composition and the lexicon. In recent years language models pre-trained on large corpora…

计算与语言 · 计算机科学 2020-03-16 Serge Sharoff

Heap's Law states that in a large enough text corpus, the number of types as a function of tokens grows as $N=KM^\beta$ for some free parameters $K,\beta$. Much has been written about how this result and various generalizations can be…

计算与语言 · 计算机科学 2019-01-04 Victor Davis

We present the QuranMorph corpus, a morphologically annotated corpus for the Quran (77,429 tokens). Each token in the QuranMorph was manually lemmatized and tagged with its part-of-speech by three expert linguists. The lemmatization process…

计算与语言 · 计算机科学 2025-06-24 Diyam Akra , Tymaa Hammouda , Mustafa Jarrar

There are different ways of measuring diversity in complex systems. In particular, in language, lexical diversity is characterized in terms of the type-token ratio and the word entropy. We here investigate both diversity metrics in six…

计算与语言 · 计算机科学 2025-07-16 Pablo Rosillo-Rodes , Maxi San Miguel , David Sanchez

Current large language models demonstrate deficiencies in understanding low-resource languages, particularly the minority languages in China. This limitation stems from the scarcity of available pre-training data. To address this…

计算与语言 · 计算机科学 2024-06-14 Chen Zhang , Mingxu Tao , Quzhe Huang , Jiuheng Lin , Zhibin Chen , Yansong Feng
‹ 上一页 1 2 3 10 下一页 ›