中文
相关论文

相关论文: Open the Data! Chuvash Datasets

200 篇论文

Spoken language datasets are vital for advancing linguistic research, Natural Language Processing, and speech technology. However, resources dedicated to Italian, a linguistically rich and diverse Romance language, remain underexplored…

计算与语言 · 计算机科学 2025-03-13 Marco Giordano , Claudia Rinaldi

Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to foster research in…

计算与语言 · 计算机科学 2020-10-27 Changhan Wang , Anne Wu , Juan Pino

Despite having a population of twenty million, Kazakhstan's culture and language remain underrepresented in the field of natural language processing. Although large language models (LLMs) continue to advance worldwide, progress in Kazakh…

Research in question answering datasets and models has gained a lot of attention in the research community. Many of them release their own question answering datasets as well as the models. There is tremendous progress that we have seen in…

计算与语言 · 计算机科学 2021-12-28 Andreas Chandra , Affandy Fahrizain , Ibrahim , Simon Willyanto Laufried

The recent advances in deep-learning have led to the development of highly sophisticated systems with an unquenchable appetite for data. On the other hand, building good deep-learning models for low-resource languages remains a challenging…

计算与语言 · 计算机科学 2024-02-20 Maithili Sabane , Onkar Litake , Aman Chadha

While natural language processing tools have been developed extensively for some of the world's languages, a significant portion of the world's over 7000 languages are still neglected. One reason for this is that evaluation datasets do not…

计算与语言 · 计算机科学 2024-06-05 Chunlan Ma , Ayyoob ImaniGooghari , Haotian Ye , Renhao Pei , Ehsaneddin Asgari , Hinrich Schütze

Data users need relevant context and research expertise to effectively search for and identify relevant datasets. Leading data providers, such as the Inter-university Consortium for Political and Social Research (ICPSR), offer standardized…

信息检索 · 计算机科学 2023-05-31 Lizhou Fan , Sara Lafia , Lingyao Li , Fangyuan Yang , Libby Hemphill

Parallel datasets are vital for performing and evaluating any kind of multilingual task. However, in the cases where one of the considered language pairs is a low-resource language, the existing top-down parallel data such as corpora are…

计算与语言 · 计算机科学 2023-09-26 Kasun Wickramasinghe , Nisansa de Silva

This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two…

音频与语音处理 · 电气工程与系统科学 2021-09-09 Saida Mussakhojayeva , Aigerim Janaliyeva , Almas Mirzakhmetov , Yerbolat Khassanov , Huseyin Atakan Varol

Speech translation for Indian languages remains a challenging task due to the scarcity of large-scale, publicly available datasets that capture the linguistic diversity and domain coverage essential for real-world applications. Existing…

We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages from different families and inherits…

计算与语言 · 计算机科学 2024-08-08 Beomseok Lee , Ioan Calapodescu , Marco Gaido , Matteo Negri , Laurent Besacier

Although numerous datasets have been developed to support dialogue systems, most existing chit-chat datasets overlook the cultural nuances inherent in natural human conversations. To address this gap, we introduce SEADialogues, a culturally…

The paper discusses the creation of a multimodal dataset of Russian-language scientific papers and testing of existing language models for the task of automatic text summarization. A feature of the dataset is its multimodal data, which…

计算与语言 · 计算机科学 2024-05-14 Alena Tsanda , Elena Bruches

Creating spoken dialogue datasets is methodologically challenging, and these challenges are amplified when the goal is to build multilingual, multi-parallel datasets at scale. This work introduces HEALTHDIAL, a large-scale, multilingual,…

计算与语言 · 计算机科学 2026-05-29 Songbo Hu , Yinhong Liu , Ej Zhou , Evgeniia Razumovskaia , Xiaobin Wang , Alexander Fraser , Ivan Vulić , Anna Korhonen

OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podcasts, talk shows, teleconferences, and other conversations.…

计算与语言 · 计算机科学 2025-09-08 Wei Chu , Yuanzhe Dong , Ke Tan , Dong Han , Xavier Menendez-Pidal , Ruchao Fan , Chenfeng Miao , Chanwoo Kim , Bhiksha Raj , Rita Singh

We present a new large-scale multilingual video description dataset, VATEX, which contains over 41,250 videos and 825,000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation…

计算机视觉与模式识别 · 计算机科学 2020-06-18 Xin Wang , Jiawei Wu , Junkun Chen , Lei Li , Yuan-Fang Wang , William Yang Wang

Question answering (QA) in English has been widely explored, but multilingual datasets are relatively new, with several methods attempting to bridge the gap between high- and low-resourced languages using data augmentation through…

计算与语言 · 计算机科学 2021-06-01 Arnab Debnath , Navid Rajabi , Fardina Fathmiul Alam , Antonios Anastasopoulos

The rising demand for inclusive speech technologies amplifies the need for multilingual datasets for Natural Language Processing (NLP) research. However, limited awareness of existing task-specific resources in low-resource languages…

计算与语言 · 计算机科学 2026-03-02 Swati Sharma , Divya V. Sharma , Anubha Gupta

In the rapidly evolving digital era, there is an increasing demand for concise information as individuals seek to distil key insights from various sources. Recent attention from researchers on Multi-document Summarisation (MDS) has resulted…

计算与语言 · 计算机科学 2024-09-19 Kushan Hewapathirana , Nisansa de Silva , C. D. Athuraliya
‹ 上一页 1 2 3 10 下一页 ›