English
Related papers

Related papers: Uzbek Cyrillic-Latin-Cyrillic Machine Transliterat…

200 papers

Abugida refers to a phonogram writing system where each syllable is represented using a single consonant or typographic ligature, along with a default vowel or optional diacritic(s) to denote other vowels. However, texting in these…

Computation and Language · Computer Science 2021-03-31 Sourav Ghosh , Sourabh Vasant Gothe , Chandramouli Sanchi , Barath Raj Kandur Raja

Recently, neural machine translation (NMT) has emerged as a powerful alternative to conventional statistical approaches. However, its performance drops considerably in the presence of morphologically rich languages (MRLs). Neural engines…

Computation and Language · Computer Science 2018-04-19 Peyman Passban , Qun Liu , Andy Way

We introduce conditional unigram tokenization, a novel approach that extends unigram tokenization by conditioning target token probabilities on source-language tokens from parallel data. Given a fixed source tokenizer, our method learns a…

Computation and Language · Computer Science 2025-07-11 Gianluca Vico , Jindřinch Libovický

Script diversity presents a challenge to Multilingual Language Models (MLLM) by reducing lexical overlap among closely related languages. Therefore, transliterating closely related languages that use different writing scripts to a common…

Computation and Language · Computer Science 2023-08-01 Ibraheem Muhammad Moosa , Mahmud Elahi Akhter , Ashfia Binte Habib

Decoder-only large language models (LLMs) excel in high-resource languages across various tasks through few-shot or even zero-shot in-context learning (ICL). However, their performance often does not transfer well to low-resource languages,…

Computation and Language · Computer Science 2024-07-03 Chunlan Ma , Yihong Liu , Haotian Ye , Hinrich Schütze

Transliteration is a key component of machine translation systems and software internationalization. This paper demonstrates that neural sequence-to-sequence models obtain state of the art or close to state of the art results on existing…

Computation and Language · Computer Science 2016-11-01 Mihaela Rosca , Thomas Breuel

The Bavarian Academy of Sciences and Humanities aims to digitize its Medieval Latin Dictionary. This dictionary entails record cards referring to lemmas in medieval Latin, a low-resource language. A crucial step of the digitization process…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Philipp Koch , Gilary Vera Nuñez , Esteban Garces Arias , Christian Heumann , Matthias Schöffel , Alexander Häberlin , Matthias Aßenmacher

Urdu is a cursive script language and has similarities with Arabic and many other South Asian languages. Urdu is difficult to classify due to its complex geometrical and morphological structure. Character classification can be processed…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Sumaiya Fazal , Sheeraz Ahmed

In this paper we present a bottom up procedure for segmentation of text lines written or printed in the Latin script. The proposed method uses a combination of image morphology, feature extraction and Gaussian mixture model to perform this…

Computer Vision and Pattern Recognition · Computer Science 2017-10-10 Himanshu Jain , Archana Praveen Kumar

This study introduces AyutthayaAlpha, an advanced transformer-based machine learning model designed for the transliteration of Thai proper names into Latin script. Our system achieves state-of-the-art performance with 82.32% first-token…

Computation and Language · Computer Science 2024-12-06 Davor Lauc , Attapol Rutherford , Weerin Wongwarawipatr

Large Language Models (LLMs) are now capable of generating text that closely resembles human writing, making them powerful tools for content creation, but this growing ability has also made it harder to tell whether a piece of text was…

Computation and Language · Computer Science 2025-10-21 Muhammad Ammar , Hadiya Murad Hadi , Usman Majeed Butt

Neural Machine Translation models have replaced the conventional phrase based statistical translation methods since the former takes a generic, scalable, data-driven approach rather than relying on manual, hand-crafted features. The neural…

Computation and Language · Computer Science 2017-12-11 Mehreen Alam , Sibt ul Hussain

Cross-lingual text classification leverages text classifiers trained in a high-resource language to perform text classification in other languages with no or minimal fine-tuning (zero/few-shots cross-lingual transfer). Nowadays,…

Computation and Language · Computer Science 2023-06-09 Inigo Jauregi Unanue , Gholamreza Haffari , Massimo Piccardi

Multilingual Large Language Models (LLMs) often provide suboptimal performance on low-resource languages like Urdu. This paper introduces UrduLLaMA 1.0, a model derived from the open-source Llama-3.1-8B-Instruct architecture and continually…

Computation and Language · Computer Science 2025-02-25 Layba Fiaz , Munief Hassan Tahir , Sana Shams , Sarmad Hussain

Pretrained language models based on the Transformer architecture have achieved state-of-the-art results in various natural language processing tasks such as part-of-speech tagging, named entity recognition, and question answering. However,…

Computation and Language · Computer Science 2021-08-24 B. Mansurov , A. Mansurov

Most human languages use scripts other than the Latin alphabet. Search users in these languages often formulate their information needs in a transliterated -- usually Latinized -- form for ease of typing. For example, Greek speakers might…

Information Retrieval · Computer Science 2025-05-14 Andreas Chari , Iadh Ounis , Sean MacAvaney

A hybrid approach for the transliteration of Algerian Arabizi: A primary study In this paper, we present a hybrid approach for the transliteration of the Algerian Arabizi. We define a set of rules enable us the passage from Arabizi to…

Computation and Language · Computer Science 2018-08-13 Imane Guellil , Faical Azouaou , Fodil Benali , Ala-Eddine Hachani , Houda Saadane

Based on an annotated multimedia corpus, television series Mar{\=a}y{\=a} 2013, we dig into the question of ''automatic standardization'' of Arabic dialects for machine translation. Here we distinguish between rule-based machine translation…

Computation and Language · Computer Science 2023-01-10 Abidrabbo Alnassan

In this work, we present several deep learning models for the automatic diacritization of Arabic text. Our models are built using two main approaches, viz. Feed-Forward Neural Network (FFNN) and Recurrent Neural Network (RNN), with several…

Computation and Language · Computer Science 2019-11-12 Ali Fadel , Ibraheem Tuffaha , Bara' Al-Jawarneh , Mahmoud Al-Ayyoub

The recognition of cursive script is regarded as a subtle task in optical character recognition due to its varied representation. Every cursive script has different nature and associated challenges. As Urdu is one of cursive language that…

Computer Vision and Pattern Recognition · Computer Science 2017-05-17 Saad Bin Ahmed , Saeeda Naz , Salahuddin Swati , Muhammad Imran Razzak