English
Related papers

Related papers: KazParC: Kazakh Parallel Corpus for Machine Transl…

200 papers

One of the most major and essential tasks in natural language processing is machine translation that is now highly dependent upon multilingual parallel corpora. Through this paper, we introduce the biggest Persian-English parallel corpus…

Computation and Language · Computer Science 2020-02-03 Omid Kashefi

Machine translation has been a major motivation of development in natural language processing. Despite the burgeoning achievements in creating more efficient machine translation systems thanks to deep learning methods, parallel corpora have…

Computation and Language · Computer Science 2020-10-06 Sina Ahmadi , Hossein Hassani , Daban Q. Jaff

Recent advances in neural machine translation (NMT) have pushed the quality of machine translation systems to the point where they are becoming widely adopted to build competitive systems. However, there is still a large number of languages…

We present an open-source speech corpus for the Kazakh language. The Kazakh speech corpus (KSC) contains around 332 hours of transcribed audio comprising over 153,000 utterances spoken by participants from different regions and age groups,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-22 Yerbolat Khassanov , Saida Mussakhojayeva , Almas Mirzakhmetov , Alen Adiyev , Mukhamet Nurpeiissov , Huseyin Atakan Varol

The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered from issues such as opaque process, difficulty of…

Computation and Language · Computer Science 2025-09-22 Qiuyang Lu , Fangjian Shen , Zhengkai Tang , Qiang Liu , Hexuan Cheng , Hui Liu , Wushao Wen

Most current machine translation models are mainly trained with parallel corpora, and their translation accuracy largely depends on the quality and quantity of the corpora. Although there are billions of parallel sentences for a few…

Computation and Language · Computer Science 2022-03-01 Makoto Morishita , Katsuki Chousa , Jun Suzuki , Masaaki Nagata

We present the IIT Bombay English-Hindi Parallel Corpus. The corpus is a compilation of parallel corpora previously available in the public domain as well as new parallel corpora we collected. The corpus contains 1.49 million parallel…

Computation and Language · Computer Science 2018-05-22 Anoop Kunchukuttan , Pratik Mehta , Pushpak Bhattacharyya

While the computational processing of Kurdish has experienced a relative increase, the machine translation of this language seems to be lacking a considerable body of scientific work. This is in part due to the lack of resources especially…

Artificial Intelligence · Computer Science 2021-06-18 Zhila Amini , Mohammad Mohammadamini , Hawre Hosseini , Mehran Mansouri , Daban Jaff

We present a new Icelandic-English parallel corpus, the Icelandic Parallel Abstracts Corpus (IPAC), composed of abstracts from student theses and dissertations. The texts were collected from the Skemman repository which keeps records of all…

Computation and Language · Computer Science 2021-08-12 Haukur Barri Símonarson , Vésteinn Snæbjarnarson

Recent machine translation algorithms mainly rely on parallel corpora. However, since the availability of parallel corpora remains limited, only some resource-rich language pairs can benefit from them. We constructed a parallel corpus for…

Computation and Language · Computer Science 2020-03-17 Makoto Morishita , Jun Suzuki , Masaaki Nagata

We release Samas\=amayik, a novel, meticulously curated, large-scale Hindi-Sanskrit corpus, comprising 92,196 parallel sentences. Unlike most data available in Sanskrit, which focuses on classical era text and poetry, this corpus aggregates…

Computation and Language · Computer Science 2026-03-26 N J Karthika , Keerthana Suryanarayanan , Jahanvi Purohit , Ganesh Ramakrishnan , Jitin Singla , Anil Kumar Gourishetty

The primary objective of our work is to build a large-scale English-Thai dataset for machine translation. We construct an English-Thai machine translation dataset with over 1 million segment pairs, curated from various sources, namely news,…

Computation and Language · Computer Science 2021-08-10 Lalita Lowphansirikul , Charin Polpanumas , Attapol T. Rutherford , Sarana Nutanong

We present a new release of the Czech-English parallel corpus CzEng 2.0 consisting of over 2 billion words (2 "gigawords") in each language. The corpus contains document-level information and is filtered with several techniques to lower the…

Computation and Language · Computer Science 2020-07-08 Tom Kocmi , Martin Popel , Ondrej Bojar

Despite having a population of twenty million, Kazakhstan's culture and language remain underrepresented in the field of natural language processing. Although large language models (LLMs) continue to advance worldwide, progress in Kazakh…

We present an expanded version of our previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In the new KazakhTTS2 corpus, the overall size has increased from 93 hours to 271 hours, the number of speakers has risen from two…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-21 Saida Mussakhojayeva , Yerbolat Khassanov , Huseyin Atakan Varol

We present a new open-source parallel corpus consisting of news articles collected from the Bianet magazine, an online newspaper that publishes Turkish news, often along with their translations in English and Kurdish. In this paper, we…

Computation and Language · Computer Science 2018-05-15 Duygu Ataman

This paper introduces a pioneering English-Azerbaijani (Arabic Script) parallel corpus, designed to bridge the technological gap in language learning and machine translation (MT) for under-resourced languages. Consisting of 548,000 parallel…

In Brazil, the governmental body responsible for overseeing and coordinating post-graduate programs, CAPES, keeps records of all theses and dissertations presented in the country. Information regarding such documents can be accessed online…

Computation and Language · Computer Science 2019-05-07 Felipe Soares , Gabrielli Harumi Yamashita , Michel Jose Anzanello

Using crowdsourcing, we collected more than 10,000 URL pairs (parallel top page pairs) of bilingual websites that contain parallel documents and created a Japanese-Chinese parallel corpus of 4.6M sentence pairs from these websites. We used…

Computation and Language · Computer Science 2024-05-16 Masaaki Nagata , Makoto Morishita , Katsuki Chousa , Norihito Yasuda

Parallel data are an important part of a reliable Statistical Machine Translation (SMT) system. The more of these data are available, the better the quality of the SMT system. However, for some language pairs such as Persian-English,…

Computation and Language · Computer Science 2019-04-02 Akbar Karimi , Ebrahim Ansari , Bahram Sadeghi Bigham
‹ Prev 1 2 3 10 Next ›