English
Related papers

Related papers: KazParC: Kazakh Parallel Corpus for Machine Transl…

200 papers

We present Samanantar, the largest publicly available parallel corpora collection for Indic languages. The collection contains a total of 49.7 million sentence pairs between English and 11 Indic languages (from two language families).…

Machine translation requires large amounts of parallel text. While such datasets are abundant in domains such as newswire, they are less accessible in the biomedical domain. Chinese and English are two of the most widely spoken languages,…

Computation and Language · Computer Science 2020-05-20 Boxiang Liu , Liang Huang

We introduce KazQAD -- a Kazakh open-domain question answering (ODQA) dataset -- that can be used in both reading comprehension and full ODQA settings, as well as for information retrieval experiments. KazQAD contains just under 6,000…

Computation and Language · Computer Science 2024-04-09 Rustem Yeshpanov , Pavel Efimov , Leonid Boytsov , Ardak Shalkarbayuli , Pavel Braslavski

The linguistic diversity of India poses significant machine translation challenges, especially for underrepresented tribal languages like Bhili, which lack high-quality linguistic resources. This paper addresses the gap by introducing…

Computation and Language · Computer Science 2025-11-04 Pooja Singh , Shashwat Bhardwaj , Vaibhav Sharma , Sandeep Kumar

Lecture transcript translation helps learners understand online courses, however, building a high-quality lecture machine translation system lacks publicly available parallel corpora. To address this, we examine a framework for parallel…

Computation and Language · Computer Science 2023-11-08 Haiyue Song , Raj Dabre , Chenhui Chu , Atsushi Fujita , Sadao Kurohashi

Multimodal neural machine translation (NMT) has become an increasingly important area of research over the years because additional modalities, such as image data, can provide more context to textual data. Furthermore, the viability of…

Computation and Language · Computer Science 2020-10-20 Andrew Merritt , Chenhui Chu , Yuki Arase

We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications published in Japan and the United States from 2000 to 2021.…

Computation and Language · Computer Science 2025-08-25 Masaaki Nagata , Katsuki Chousa , Norihito Yasuda

We present ASCAT (Arabic Scientific Corpus for Advanced Translation), a high-quality English-Arabic parallel benchmark corpus designed for scientific translation evaluation constructed through a systematic multi-engine translation and human…

Computation and Language · Computer Science 2026-04-02 Serry Sibaee , Khloud Al Jallad , Zineb Yousfi , Israa Elsayed Elhosiny , Yousra El-Ghawi , Batool Balah , Omer Nacar

This paper describes the acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. It will be helpful for machine translation of a low-resource language, Amharic. We freely released the corpus for…

Computation and Language · Computer Science 2022-05-04 Andargachew Mekonnen Gezmu , Andreas Nürnberger , Tesfaye Bayu Bati

Most legal text in the Indian judiciary is written in complex English due to historical reasons. However, only a small fraction of the Indian population is comfortable in reading English. Hence legal text needs to be made available in…

Computation and Language · Computer Science 2024-11-08 Sayan Mahapatra , Debtanu Datta , Shubham Soni , Adrijit Goswami , Saptarshi Ghosh

Although the parallel corpus has an irreplaceable role in machine translation, its scale and coverage is still beyond the actual needs. Non-parallel corpus resources on the web have an inestimable potential value in machine translation and…

Computation and Language · Computer Science 2014-05-23 Lijiang Chen

We propose a method for efficiently finding all parallel passages in a large corpus, even if the passages are not quite identical due to rephrasing and orthographic variation. The key ideas are the representation of each word in the corpus…

Computation and Language · Computer Science 2023-06-22 Avi Shmidman , Moshe Koppel , Ely Porat

This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-09 Saida Mussakhojayeva , Aigerim Janaliyeva , Almas Mirzakhmetov , Yerbolat Khassanov , Huseyin Atakan Varol

The data article presents the large bilingual parallel corpus of low-resourced language pair Sanskrit-Hindi, named SAHAAYAK 2023. The corpus contains total of 1.5M sentence pairs between Sanskrit and Hindi. To make the universal usability…

Computation and Language · Computer Science 2023-07-04 Vishvajitsinh Bakrola , Jitendra Nasariwala

This paper introduces PEACH, a sentence-aligned parallel English-Arabic corpus of healthcare texts encompassing patient information leaflets and educational materials. The corpus contains 51,671 parallel sentences, totaling approximately…

Computation and Language · Computer Science 2025-08-11 Rania Al-Sabbagh

This paper presents KazSAnDRA, a dataset developed for Kazakh sentiment analysis that is the first and largest publicly available dataset of its kind. KazSAnDRA comprises an extensive collection of 180,064 reviews obtained from various…

Computation and Language · Computer Science 2024-04-11 Rustem Yeshpanov , Huseyin Atakan Varol

Neural Machine Translation with its significant results, still has a great problem: lack or absence of parallel corpus for many languages. This article suggests a method for generating considerable amount of parallel corpus for any language…

Computation and Language · Computer Science 2018-04-12 Farshad Jafari

Parallel text is required for building high-quality machine translation (MT) systems, as well as for other multilingual NLP applications. For many South Asian languages, such data is in short supply. In this paper, we described a new…

Computation and Language · Computer Science 2020-01-28 Barry Haddow , Faheem Kirefu

Analogy-making is central to human cognition, allowing us to adapt to novel situations -- an ability that current AI systems still lack. Most analogy datasets today focus on simple analogies (e.g., word analogies); datasets including…

Computation and Language · Computer Science 2024-05-15 Oren Sultan , Yonatan Bitton , Ron Yosef , Dafna Shahaf

Code-mixing is the phenomenon of using more than one language in a sentence. It is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to…

Computation and Language · Computer Science 2020-04-21 Vivek Srivastava , Mayank Singh