English
Related papers

Related papers: A Rule-based Kurdish Text Transliteration System

200 papers

Background: The accuracy of spelling in Electronic Health Records (EHRs) is a critical factor for efficient clinical care, research, and ensuring patient safety. The Persian language, with its abundant vocabulary and complex…

Computation and Language · Computer Science 2024-08-08 Seyed Mohammad Sadegh Dashti , Seyedeh Fatemeh Dashti

Character recognition techniques for printed documents are widely used for English language. However, the systems that are implemented to recognize Asian languages struggle to increase the accuracy of recognition. Among other Asian…

Computer Vision and Pattern Recognition · Computer Science 2014-12-25 G. I. Gunarathna , M. A. P. Chamikara , R. G. Ragel

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Computation and Language · Computer Science 2025-11-10 Haneen Al-Homoud , Asma Ibrahim , Murtadha Al-Jubran , Fahad Al-Otaibi , Yazeed Al-Harbi , Daulet Toibazar , Kesen Wang , Pedro J. Moreno

Recognition of Arabic characters is essential for natural language processing and computer vision fields. The need to recognize and classify the handwritten Arabic letters and characters are essentially required. In this paper, we present…

Computer Vision and Pattern Recognition · Computer Science 2020-09-29 Mahmoud Shams , Amira. A. Elsonbaty , Wael. Z. ElSawy

We propose a novel multitask learning method for diacritization which trains a model to both diacritize and translate. Our method addresses data sparsity by exploiting large, readily available bitext corpora. Furthermore, translation…

Computation and Language · Computer Science 2021-09-30 Brian Thompson , Ali Alshehri

This paper reports on the preliminary phase of our ongoing research towards developing an intelligent tutoring environment for Turkish grammar. One of the components of this environment is a corpus search tool which, among other aspects of…

cmp-lg · Computer Science 2016-08-31 H. Altay Guvenir , Kemal Oflazer

Online Arabic cursive character recognition is still a big challenge due to the existing complexities including Arabic cursive script styles, writing speed, writer mood and so forth. Due to these unavoidable constraints, the accuracy of…

Computer Vision and Pattern Recognition · Computer Science 2021-01-18 Amjad Rehman

Word segmentation plays a pivotal role in improving any Arabic NLP application. Therefore, a lot of research has been spent in improving its accuracy. Off-the-shelf tools, however, are: i) complicated to use and ii) domain/dialect…

Computation and Language · Computer Science 2017-09-05 Hassan Sajjad , Fahim Dalvi , Nadir Durrani , Ahmed Abdelali , Yonatan Belinkov , Stephan Vogel

State-of-the-art Natural Language Processing algorithms rely heavily on efficient word segmentation. Urdu is amongst languages for which word segmentation is a complex task as it exhibits space omission as well as space insertion issues.…

Computation and Language · Computer Science 2018-06-15 Haris Bin Zia , Agha Ali Raza , Awais Athar

In order to provide benchmark performance for Urdu text document classification, the contribution of this paper is manifold. First, it pro-vides a publicly available benchmark dataset manually tagged against 6 classes. Second, it…

Computation and Language · Computer Science 2020-03-04 Muhammad Nabeel Asim , Muhammad Usman Ghani , Muhammad Ali Ibrahim , Sheraz Ahmad , Waqar Mahmood , Andreas Dengel

A large amount of feedback was collected over the years. Many feedback analysis models have been developed focusing on the English language. Recognizing the concept of feedback is challenging and crucial in languages which do not have…

Computation and Language · Computer Science 2023-02-24 Zolzaya Dashdorj , Tsetsentsengel Munkhbayar , Stanislav Grigorev

Machine translation is research based area where evaluation is very important phenomenon for checking the quality of MT output. The work is based on the evaluation of English to Urdu Machine translation. In this research work we have…

Computation and Language · Computer Science 2013-10-03 Vaishali Gupta , Nisheeth Joshi , Iti Mathur

In this study, we develop and assess new corpus selection and training methodologies to improve the effectiveness of Turkish language models. Specifically, we adapted Large Language Model generated datasets and translated English datasets…

Computation and Language · Computer Science 2024-12-05 H. Toprak Kesgin , M. Kaan Yuce , Eren Dogan , M. Egemen Uzun , Atahan Uz , Elif Ince , Yusuf Erdem , Osama Shbib , Ahmed Zeer , M. Fatih Amasyali

Handwritten text recognition is an active research area in the field of deep learning and artificial intelligence to convert handwritten text into machine-understandable. A lot of work has been done for other languages, especially for…

Computer Vision and Pattern Recognition · Computer Science 2021-03-10 Muhammad Kashif

The ability to synthesize spoken language from text has greatly facilitated access to digital content with the advances in text-to-speech technology. However, effective TTS development for low-resource languages, such as Central Kurdish…

Computation and Language · Computer Science 2024-09-25 Abdulhady Abas Abdullah , Sabat Salih Muhamad , Hadi Veisi

Ironic identification is a challenging task in Natural Language Processing, particularly when dealing with languages that differ in syntax and cultural context. In this work, we aim to detect irony in Urdu by translating an English Ironic…

Computation and Language · Computer Science 2025-10-28 Fiaz Ahmad , Nisar Hussain , Amna Qasim , Momina Hafeez , Muhammad Usman Grigori Sidorov , Alexander Gelbukh

Spelling correction is the task of identifying spelling mistakes, typos, and grammatical mistakes in a given text and correcting them according to their context and grammatical structure. This work introduces "AraSpell," a framework for…

Computation and Language · Computer Science 2024-05-14 Mahmoud Salhab , Faisal Abu-Khzam

Kazakh is a Turkic language using the Arabic, Cyrillic, and Latin scripts, making it unique in terms of optical character recognition (OCR). Work on OCR for low-resource Kazakh scripts is very scarce, and no OCR benchmarks or images exist…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Henry Gagnier , Sophie Gagnier , Ashwin Kirubakaran

Machine Transliteration has come out to be an emerging and a very important research area in the field of machine translation. Transliteration basically aims to preserve the phonological structure of words. Proper transliteration of name…

Computation and Language · Computer Science 2013-07-17 Deepti Bhalla , Nisheeth Joshi , Iti Mathur

Text classification has seen an increased use in both academic and industry settings. Though rule based methods have been fairly successful, supervised machine learning has been shown to be most successful for most languages, where most…

Computation and Language · Computer Science 2021-04-26 Deniz Kavi
‹ Prev 1 3 4 5 6 7 10 Next ›