English
Related papers

Related papers: KazParC: Kazakh Parallel Corpus for Machine Transl…

200 papers

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source…

Computation and Language · Computer Science 2021-06-16 Elizabeth Salesky , Matthew Wiesner , Jacob Bremerman , Roldano Cattoni , Matteo Negri , Marco Turchi , Douglas W. Oard , Matt Post

Automated documentation of programming source code and automated code generation from natural language are challenging tasks of both practical and scientific interest. Progress in these areas has been limited by the low availability of…

Computation and Language · Computer Science 2017-07-10 Antonio Valerio Miceli Barone , Rico Sennrich

Objective: Today's neural machine translation (NMT) can achieve near human-level translation quality and greatly facilitates international communications, but the lack of parallel corpora poses a key problem to the development of…

Computation and Language · Computer Science 2022-02-08 Shengxuan Luo , Huaiyuan Ying , Jiao Li , Sheng Yu

We present a new, unique and freely available parallel corpus containing European Union (EU) documents of mostly legal nature. It is available in all 20 official EUanguages, with additional documents being available in the languages of the…

Computation and Language · Computer Science 2007-05-23 Ralf Steinberger , Bruno Pouliquen , Anna Widiger , Camelia Ignat , Tomaz Erjavec , Dan Tufis , Daniel Varga

We introduce S2ORC, a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for…

Computation and Language · Computer Science 2020-07-08 Kyle Lo , Lucy Lu Wang , Mark Neumann , Rodney Kinney , Dan S. Weld

Ancient Buddhist literature features frequent, yet often unannotated, textual parallels spread across diverse languages: Sanskrit, P\=ali, Buddhist Chinese, Tibetan, and more. The scale of this material makes manual examination prohibitive.…

Computation and Language · Computer Science 2026-01-13 Sebastian Nehrdich , Kurt Keutzer

We describe PARANMT-50M, a dataset of more than 50 million English-English sentential paraphrase pairs. We generated the pairs automatically by using neural machine translation to translate the non-English side of a large parallel corpus,…

Computation and Language · Computer Science 2018-04-23 John Wieting , Kevin Gimpel

This paper presents the first benchmark for the task of automatic part-of-speech (POS) tagging for the Tajik language. Despite the existence of multilingual language models demonstrating high effectiveness for many of the world's languages,…

Computation and Language · Computer Science 2026-05-07 Mullosharaf K. Arabov

We present a parallel machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. We used a transformer-based translator to generate initial translations in Akuapem Twi, which were later verified and corrected…

This paper presents BSTC (Baidu Speech Translation Corpus), a large-scale Chinese-English speech translation dataset. This dataset is constructed based on a collection of licensed videos of talks or lectures, including about 68 hours of…

Computation and Language · Computer Science 2021-04-28 Ruiqing Zhang , Xiyang Wang , Chuanqiang Zhang , Zhongjun He , Hua Wu , Zhi Li , Haifeng Wang , Ying Chen , Qinfei Li

The increasing volume of scientific research necessitates effective communication across language barriers. Machine translation (MT) offers a promising solution for accessing international publications. However, the scientific domain…

Computation and Language · Computer Science 2026-05-21 Dimitris Roussis , Sokratis Sofianopoulos , Stelios Piperidis

Translating culture-related content is vital for effective cross-cultural communication. However, many culture-specific items (CSIs) often lack viable translations across languages, making it challenging to collect high-quality, diverse…

Computation and Language · Computer Science 2024-10-22 Binwei Yao , Ming Jiang , Tara Bobinac , Diyi Yang , Junjie Hu

In this paper, we present a corpus for use in automatic readability assessment and automatic text simplification of German. The corpus is compiled from web sources and consists of approximately 211,000 sentences. As a novel contribution, it…

Computation and Language · Computer Science 2019-09-20 Alessia Battisti , Sarah Ebling

This paper presents the NICT's participation in the WMT18 shared parallel corpus filtering task. The organizers provided 1 billion words German-English corpus crawled from the web as part of the Paracrawl project. This corpus is too noisy…

Computation and Language · Computer Science 2018-10-15 Rui Wang , Benjamin Marie , Masao Utiyama , Eiichiro Sumita

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real online one-to-one…

Computation and Language · Computer Science 2022-06-28 Chengfei Li , Shuhao Deng , Yaoping Wang , Guangjing Wang , Yaguang Gong , Changbin Chen , Jinfeng Bai

This study presents several contributions for the Karakalpak language: a FLORES+ devtest dataset translated to Karakalpak, parallel corpora for Uzbek-Karakalpak, Russian-Karakalpak and English-Karakalpak of 100,000 pairs each and…

Computation and Language · Computer Science 2024-09-09 Mukhammadsaid Mamasaidov , Abror Shopulatov

Although there are increasing and significant ties between China and Portuguese-speaking countries, there is not much parallel corpora in the Chinese-Portuguese language pair. Both languages are very populous, with 1.2 billion native…

Computation and Language · Computer Science 2018-04-06 Siyou Liu , Longyue Wang , Chao-Hong Liu

We introduce MTet, the largest publicly available parallel corpus for English-Vietnamese translation. MTet consists of 4.2M high-quality training sentence pairs and a multi-domain test set refined by the Vietnamese research community.…

Computation and Language · Computer Science 2022-10-20 Chinh Ngo , Trieu H. Trinh , Long Phan , Hieu Tran , Tai Dang , Hieu Nguyen , Minh Nguyen , Minh-Thang Luong

This paper presents the first comprehensive comparative analysis of modern machine learning architectures for transliteration between Tajik (Cyrillic script) and Persian (Arabic script). A key contribution is the creation and validation of…

Computation and Language · Computer Science 2026-05-05 Mullosharaf K. Arabov

This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 69,441 sentences spanning 1+ million words, carefully curated to cover…

Computation and Language · Computer Science 2025-06-17 Khalid N. Elmadani , Nizar Habash , Hanada Taha-Thomure
‹ Prev 1 4 5 6 7 8 10 Next ›