中文
相关论文

相关论文: A Parallel Corpus of Theses and Dissertations Abst…

200 篇论文

Neural Machine Translation (NMT) has been proven to achieve impressive results. The NMT system translation results depend strongly on the size and quality of parallel corpora. Nevertheless, for many language pairs, no rich-resource parallel…

计算与语言 · 计算机科学 2019-05-23 Jinyi Zhang , Tadahiro Matsumoto

With a large amount of parallel data, neural machine translation systems are able to deliver human-level performance for sentence-level translation. However, it is costly to label a large amount of parallel data by humans. In contrast,…

计算与语言 · 计算机科学 2020-09-21 Guokun Lai , Zihang Dai , Yiming Yang

This paper presents the first publicly available version of the Carolina Corpus and discusses its future directions. Carolina is a large open corpus of Brazilian Portuguese texts under construction using web-as-corpus methodology enhanced…

Machine translation (MT) has almost achieved human parity at sentence-level translation. In response, the MT community has, in part, shifted its focus to document-level translation. However, the development of document-level MT systems is…

计算与语言 · 计算机科学 2022-10-27 Yuchen Eleanor Jiang , Tianyu Liu , Shuming Ma , Dongdong Zhang , Mrinmaya Sachan , Ryan Cotterell

Parallel corpus is a critical resource in machine learning-based translation. The task of collecting, extracting, and aligning texts in order to build an acceptable corpus for doing the translation is very tedious most especially for…

计算与语言 · 计算机科学 2021-11-01 Kristine Mae Adlaon , Nelson Marcos

In this paper, we present a parallel Spanish-Mazatec and Spanish-Mixtec corpus for machine translation (MT) tasks, where Mazatec and Mixtec are two indigenous Mexican languages. We evaluated the usability of the collected corpus using three…

Several recent papers claim human parity at sentence-level Machine Translation (MT), especially in high-resource languages. Thus, in response, the MT community has, in part, shifted its focus to document-level translation. Translating…

计算与语言 · 计算机科学 2023-05-19 Yuchen Eleanor Jiang , Tianyu Liu , Shuming Ma , Dongdong Zhang , Mrinmaya Sachan , Ryan Cotterell

Recent works have shown that synthetic parallel data automatically generated by translation models can be effective for various neural machine translation (NMT) issues. In this study, we build NMT systems using only synthetic parallel data.…

计算与语言 · 计算机科学 2017-09-19 Jaehong Park , Jongyoon Song , Sungroh Yoon

Major advancement in the performance of machine translation models has been made possible in part thanks to the availability of large-scale parallel corpora. But for most languages in the world, the existence of such corpora is rare.…

计算与语言 · 计算机科学 2021-04-14 Felermino D. M. A. Ali , Andrew Caines , Jaimito L. A. Malavi

Significant advances have been made in natural language processing in recent years. However, our current deep learning approach to language modeling requires substantial resources in terms of data and computation. One of the side effects of…

计算与语言 · 计算机科学 2025-07-25 Nicholas Kluge Corrêa , Aniket Sen , Sophia Falk , Shiza Fatimah

This paper describes the machine translation systems developed by the Universidade Federal do Rio Grande do Sul (UFRGS) team for the biomedical translation shared task. Our systems are based on statistical machine translation and neural…

计算与语言 · 计算机科学 2019-05-07 Felipe Soares , Karin Becker

Document-level neural machine translation (NMT) has outperformed sentence-level NMT on a number of datasets. However, document-level NMT is still not widely adopted in real-world translation systems mainly due to the lack of large-scale…

计算与语言 · 计算机科学 2023-04-21 Yusser Al Ghussin , Jingyi Zhang , Josef van Genabith

Text alignment and text quality are critical to the accuracy of Machine Translation (MT) systems, some NLP tools, and any other text processing tasks requiring bilingual data. This research proposes a language independent bi-sentence…

计算与语言 · 计算机科学 2015-10-16 Krzysztof Wołk

Parallel text is required for building high-quality machine translation (MT) systems, as well as for other multilingual NLP applications. For many South Asian languages, such data is in short supply. In this paper, we described a new…

计算与语言 · 计算机科学 2020-01-28 Barry Haddow , Faheem Kirefu

Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limits its ability to…

计算与语言 · 计算机科学 2025-10-22 Yingli Shen , Wen Lai , Shuo Wang , Ge Gao , Kangyang Luo , Alexander Fraser , Maosong Sun

The primary objective of our work is to build a large-scale English-Thai dataset for machine translation. We construct an English-Thai machine translation dataset with over 1 million segment pairs, curated from various sources, namely news,…

计算与语言 · 计算机科学 2021-08-10 Lalita Lowphansirikul , Charin Polpanumas , Attapol T. Rutherford , Sarana Nutanong

In this paper, a novel approach is proposed to automatically construct parallel discourse corpus for dialogue machine translation. Firstly, the parallel subtitle data and its corresponding monolingual movie script data are crawled and…

计算与语言 · 计算机科学 2016-05-24 Longyue Wang , Xiaojun Zhang , Zhaopeng Tu , Andy Way , Qun Liu

Automatic Essay Scoring (AES) is defined as the computer technology that evaluates and scores the written essays, aiming to provide computational models to grade essays either automatically or with minimal human involvement. While there are…

计算与语言 · 计算机科学 2021-05-20 Jeziel C. Marinho , Rafael T. Anchieta , Raimundo S. Moura

This paper presents the NICT's participation in the WMT18 shared parallel corpus filtering task. The organizers provided 1 billion words German-English corpus crawled from the web as part of the Paracrawl project. This corpus is too noisy…

计算与语言 · 计算机科学 2018-10-15 Rui Wang , Benjamin Marie , Masao Utiyama , Eiichiro Sumita

We present Samanantar, the largest publicly available parallel corpora collection for Indic languages. The collection contains a total of 49.7 million sentence pairs between English and 11 Indic languages (from two language families).…