English
Related papers

Related papers: Uzbek Cyrillic-Latin-Cyrillic Machine Transliterat…

200 papers

Although machine translation systems are mostly designed to serve in the general domain, there is a growing tendency to adapt these systems to other domains like literary translation. In this paper, we focus on English-Turkish literary…

Computation and Language · Computer Science 2023-07-24 Zeynep Yirmibeşoğlu , Olgun Dursun , Harun Dallı , Mehmet Şahin , Ena Hodzik , Sabri Gürses , Tunga Güngör

We report upon the results of a research and prototype building project \emph{Worldly~OCR} dedicated to developing new, more accurate image-to-text conversion software for several languages and writing systems. These include the cursive…

Computer Vision and Pattern Recognition · Computer Science 2020-05-19 Marek Rychlik , Dwight Nwaigwe , Yan Han , Dylan Murphy

Text classification has become a crucial task in various fields, leading to a significant amount of research on developing automated text classification systems for national and international languages. However, there is a growing need for…

Computation and Language · Computer Science 2023-05-08 Mursal Dawodi , Jawid Ahmad Baktash

This paper advances NLP research for the low-resource Uzbek language by evaluating two previously untested monolingual Uzbek BERT models on the part-of-speech (POS) tagging task and introducing the first publicly available UPOS-tagged…

Computation and Language · Computer Science 2025-01-20 Latofat Bobojonova , Arofat Akhundjanova , Phil Ostheimer , Sophie Fellenz

In this paper we propose a novel neural approach for automatic decipherment of lost languages. To compensate for the lack of strong supervision signal, our model design is informed by patterns in language change documented in historical…

Computation and Language · Computer Science 2019-06-18 Jiaming Luo , Yuan Cao , Regina Barzilay

Judeo-Arabic refers to Arabic variants historically spoken by Jewish communities across the Arab world, primarily during the Middle Ages. Unlike standard Arabic, it is written in Hebrew script by Jewish writers and for Jewish audiences.…

Computation and Language · Computer Science 2026-01-30 Juan Moreno Gonzalez , Bashar Alhafni , Nizar Habash

Discriminating between closely-related language varieties is considered a challenging and important task. This paper describes our submission to the DSL 2016 shared-task, which included two sub-tasks: one on discriminating similar languages…

Computation and Language · Computer Science 2016-09-27 Yonatan Belinkov , James Glass

Kurdish is written in different scripts. The two most popular scripts are Latin and Persian-Arabic. However, not all Kurdish readers are familiar with both mentioned scripts that could be resolved by automatic transliterators. So far, the…

Computation and Language · Computer Science 2021-10-26 Hossein Hassani

In order to provide benchmark performance for Urdu text document classification, the contribution of this paper is manifold. First, it pro-vides a publicly available benchmark dataset manually tagged against 6 classes. Second, it…

Computation and Language · Computer Science 2020-03-04 Muhammad Nabeel Asim , Muhammad Usman Ghani , Muhammad Ali Ibrahim , Sheraz Ahmad , Waqar Mahmood , Andreas Dengel

Despite notable advances in large language models (LLMs), reliable evaluation of text generation tasks such as text style transfer (TST) remains an open challenge. Existing research has shown that automatic metrics often correlate poorly…

Computation and Language · Computer Science 2026-03-05 Vitaly Protasov , Nikolay Babakov , Daryna Dementieva , Alexander Panchenko

We present Latin BERT, a contextual language model for the Latin language, trained on 642.7 million words from a variety of sources spanning the Classical era to the 21st century. In a series of case studies, we illustrate the affordances…

Computation and Language · Computer Science 2020-09-22 David Bamman , Patrick J. Burns

We present the second ever evaluated Arabic dialect-to-dialect machine translation effort, and the first to leverage external resources beyond a small parallel corpus. The subject has not previously received serious attention due to lack of…

Computation and Language · Computer Science 2017-12-19 Alexander Erdmann , Nizar Habash , Dima Taji , Houda Bouamor

This work presents a morphological analyzer for the Uzbek language using a finite state machine. The proposed methodology is a morphologic analysis of Uzbek words by using an affix striping to find a root and without including any lexicon.…

Computation and Language · Computer Science 2022-05-23 Maksud Sharipov , Ulugbek Salaev

Neural machine translation (NMT) has achieved notable performance recently. However, this approach has not been widely applied to the translation task between Chinese and Uyghur, partly due to the limited parallel data resource and the…

Computation and Language · Computer Science 2017-06-28 Shiyue Zhang , Gulnigar Mahmut , Dong Wang , Askar Hamdulla

As large language models (LLMs) are trained on increasingly diverse and extensive multilingual corpora, they demonstrate cross-lingual transfer capabilities. However, these capabilities often fail to effectively extend to low-resource…

Computation and Language · Computer Science 2025-09-23 Wenhao Zhuang , Yuan Sun , Xiaobing Zhao

Word segmentation plays a pivotal role in improving any Arabic NLP application. Therefore, a lot of research has been spent in improving its accuracy. Off-the-shelf tools, however, are: i) complicated to use and ii) domain/dialect…

Computation and Language · Computer Science 2017-09-05 Hassan Sajjad , Fahim Dalvi , Nadir Durrani , Ahmed Abdelali , Yonatan Belinkov , Stephan Vogel

The Latin script is often used to informally write languages with non-Latin native scripts. In many cases (e.g., most languages in India), the lack of conventional spelling in the Latin script results in high spelling variability. Such…

Computation and Language · Computer Science 2025-11-19 Adrian Benton , Alexander Gutkin , Christo Kirov , Brian Roark

This paper presents a novel semantic-based phrase translation model. A pair of source and target phrases are projected into continuous-valued vector representations in a low-dimensional latent semantic space, where their translation score…

Computation and Language · Computer Science 2013-12-03 Jianfeng Gao , Xiaodong He , Wen-tau Yih , Li Deng

A prototype system for the transliteration of diacritics-less Arabic manuscripts at the sub-word or part of Arabic word (PAW) level is developed. The system is able to read sub-words of the input manuscript using a set of skeleton-based…

Computer Vision and Pattern Recognition · Computer Science 2013-06-27 Reza Farrahi Moghaddam , Mohamed Cheriet , Thomas Milo , Robert Wisnovsky

Over recent years a lot of research papers and studies have been published on the development of effective approaches that benefit from a large amount of user-generated content and build intelligent predictive models on top of them. This…

Computation and Language · Computer Science 2021-01-21 Mohammad Kasra Habib