中文
相关论文

相关论文: KazParC: Kazakh Parallel Corpus for Machine Transl…

200 篇论文

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source…

Automated documentation of programming source code and automated code generation from natural language are challenging tasks of both practical and scientific interest. Progress in these areas has been limited by the low availability of…

计算与语言 · 计算机科学 2017-07-10 Antonio Valerio Miceli Barone , Rico Sennrich

Objective: Today's neural machine translation (NMT) can achieve near human-level translation quality and greatly facilitates international communications, but the lack of parallel corpora poses a key problem to the development of…

计算与语言 · 计算机科学 2022-02-08 Shengxuan Luo , Huaiyuan Ying , Jiao Li , Sheng Yu

We present a new, unique and freely available parallel corpus containing European Union (EU) documents of mostly legal nature. It is available in all 20 official EUanguages, with additional documents being available in the languages of the…

计算与语言 · 计算机科学 2007-05-23 Ralf Steinberger , Bruno Pouliquen , Anna Widiger , Camelia Ignat , Tomaz Erjavec , Dan Tufis , Daniel Varga

We introduce S2ORC, a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for…

计算与语言 · 计算机科学 2020-07-08 Kyle Lo , Lucy Lu Wang , Mark Neumann , Rodney Kinney , Dan S. Weld

Ancient Buddhist literature features frequent, yet often unannotated, textual parallels spread across diverse languages: Sanskrit, P\=ali, Buddhist Chinese, Tibetan, and more. The scale of this material makes manual examination prohibitive.…

计算与语言 · 计算机科学 2026-01-13 Sebastian Nehrdich , Kurt Keutzer

We describe PARANMT-50M, a dataset of more than 50 million English-English sentential paraphrase pairs. We generated the pairs automatically by using neural machine translation to translate the non-English side of a large parallel corpus,…

计算与语言 · 计算机科学 2018-04-23 John Wieting , Kevin Gimpel

This paper presents the first benchmark for the task of automatic part-of-speech (POS) tagging for the Tajik language. Despite the existence of multilingual language models demonstrating high effectiveness for many of the world's languages,…

计算与语言 · 计算机科学 2026-05-07 Mullosharaf K. Arabov

We present a parallel machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. We used a transformer-based translator to generate initial translations in Akuapem Twi, which were later verified and corrected…

This paper presents BSTC (Baidu Speech Translation Corpus), a large-scale Chinese-English speech translation dataset. This dataset is constructed based on a collection of licensed videos of talks or lectures, including about 68 hours of…

计算与语言 · 计算机科学 2021-04-28 Ruiqing Zhang , Xiyang Wang , Chuanqiang Zhang , Zhongjun He , Hua Wu , Zhi Li , Haifeng Wang , Ying Chen , Qinfei Li

The increasing volume of scientific research necessitates effective communication across language barriers. Machine translation (MT) offers a promising solution for accessing international publications. However, the scientific domain…

计算与语言 · 计算机科学 2026-05-21 Dimitris Roussis , Sokratis Sofianopoulos , Stelios Piperidis

Translating culture-related content is vital for effective cross-cultural communication. However, many culture-specific items (CSIs) often lack viable translations across languages, making it challenging to collect high-quality, diverse…

计算与语言 · 计算机科学 2024-10-22 Binwei Yao , Ming Jiang , Tara Bobinac , Diyi Yang , Junjie Hu

In this paper, we present a corpus for use in automatic readability assessment and automatic text simplification of German. The corpus is compiled from web sources and consists of approximately 211,000 sentences. As a novel contribution, it…

计算与语言 · 计算机科学 2019-09-20 Alessia Battisti , Sarah Ebling

This paper presents the NICT's participation in the WMT18 shared parallel corpus filtering task. The organizers provided 1 billion words German-English corpus crawled from the web as part of the Paracrawl project. This corpus is too noisy…

计算与语言 · 计算机科学 2018-10-15 Rui Wang , Benjamin Marie , Masao Utiyama , Eiichiro Sumita

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one…

计算与语言 · 计算机科学 2022-06-28 Chengfei Li , Shuhao Deng , Yaoping Wang , Guangjing Wang , Yaguang Gong , Changbin Chen , Jinfeng Bai

This study presents several contributions for the Karakalpak language: a FLORES+ devtest dataset translated to Karakalpak, parallel corpora for Uzbek-Karakalpak, Russian-Karakalpak and English-Karakalpak of 100,000 pairs each and…

计算与语言 · 计算机科学 2024-09-09 Mukhammadsaid Mamasaidov , Abror Shopulatov

Although there are increasing and significant ties between China and Portuguese-speaking countries, there is not much parallel corpora in the Chinese-Portuguese language pair. Both languages are very populous, with 1.2 billion native…

计算与语言 · 计算机科学 2018-04-06 Siyou Liu , Longyue Wang , Chao-Hong Liu

We introduce MTet, the largest publicly available parallel corpus for English-Vietnamese translation. MTet consists of 4.2M high-quality training sentence pairs and a multi-domain test set refined by the Vietnamese research community.…

计算与语言 · 计算机科学 2022-10-20 Chinh Ngo , Trieu H. Trinh , Long Phan , Hieu Tran , Tai Dang , Hieu Nguyen , Minh Nguyen , Minh-Thang Luong

This paper presents the first comprehensive comparative analysis of modern machine learning architectures for transliteration between Tajik (Cyrillic script) and Persian (Arabic script). A key contribution is the creation and validation of…

计算与语言 · 计算机科学 2026-05-05 Mullosharaf K. Arabov

This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 69,441 sentences spanning 1+ million words, carefully curated to cover…

计算与语言 · 计算机科学 2025-06-17 Khalid N. Elmadani , Nizar Habash , Hanada Taha-Thomure