English
Related papers

Related papers: A Clustering Framework for Lexical Normalization o…

200 papers

This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. The dataset includes, for each language: 1) native script Wikipedia text; 2) a romanization…

Computation and Language · Computer Science 2020-07-03 Brian Roark , Lawrence Wolf-Sonkin , Christo Kirov , Sabrina J. Mielke , Cibu Johny , Isin Demirsahin , Keith Hall

Deep learning frameworks allowed for a remarkable advancement in semantic segmentation, but the data hungry nature of convolutional networks has rapidly raised the demand for adaptation techniques able to transfer learned knowledge from…

Computer Vision and Pattern Recognition · Computer Science 2020-11-26 Marco Toldo , Umberto Michieli , Pietro Zanuttigh

We introduce romanization encoding for script-heavy languages to optimize multilingual and code-switching Automatic Speech Recognition (ASR) systems. By adopting romanization encoding alongside a balanced concatenated tokenizer within a…

Computation and Language · Computer Science 2024-12-18 Wen Ding , Fei Jia , Hainan Xu , Yu Xi , Junjie Lai , Boris Ginsburg

Arabic is one of the languages that present special challenges to Optical character recognition (OCR). The main challenge in Arabic is that it is mostly cursive. Therefore, a segmentation process must be carried out to determine where the…

Computer Vision and Pattern Recognition · Computer Science 2017-07-05 Mahmoud A. A. Mousa , Mohammed S. Sayed , Mahmoud I. Abdalla

Semitic languages can be highly ambiguous, having several interpretations of the same surface forms, and morphologically rich, having many morphemes that realize several morphological features. This is further exacerbated for dialectal…

Computation and Language · Computer Science 2019-10-08 Nasser Zalmout , Nizar Habash

The present paper aims at presenting a lemmatization and a word-level error correction system for Sorani Kurdish. We propose a hybrid approach based on the morphological rules and a n-gram language model. We have called our lemmatization…

Computation and Language · Computer Science 2018-10-01 Shahin Salavati , Sina Ahmadi

Large Language Models (LLMs) are increasingly deployed in high-stakes clinical applications in India. Speakers of Indian languages frequently communicate using romanized text rather than native scripts, yet existing research rarely…

Computation and Language · Computer Science 2026-04-01 Manurag Khullar , Utkarsh Desai , Poorva Malviya , Aman Dalmia , Zheyuan Ryan Shi

Machine translation is research based area where evaluation is very important phenomenon for checking the quality of MT output. The work is based on the evaluation of English to Urdu Machine translation. In this research work we have…

Computation and Language · Computer Science 2013-10-03 Vaishali Gupta , Nisheeth Joshi , Iti Mathur

The understanding of the human language is quantified by identifying intents and entities. Even though classification methods that rely on labeled information are often used for the comprehension of language understanding, it is incredibly…

Computation and Language · Computer Science 2021-02-23 Johar Shabbir , Muhammad Umair Arshad , Waseem Shahzad

Oracle Bone Inscriptions (OBIs), play a crucial role in understanding ancient Chinese civilization. The automated detection of OBIs from rubbing images represents a fundamental yet challenging task in digital archaeology, primarily due to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Ye Tao , Xinran Fu , Honglin Pang , Xi Yang , Chuntao Li

Most existing OCR methods focus on alphanumeric characters due to the popularity of English and numbers, as well as their corresponding datasets. On extending the characters to more languages, recent methods have shown that training…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Jing Huang , Kevin J Liang , Rama Kovvuri , Tal Hassner

OCR character segmentation for multilingual printed documents is difficult due to the diversity of different linguistic characters. Previous approaches mainly focus on monolingual texts and are not suitable for multilingual-lingual cases.…

Computer Vision and Pattern Recognition · Computer Science 2016-11-17 Huabin Zheng , Jingyu Wang , Zhengjie Huang , Yang Yang , Rong Pan

In this work, we present the development of a reverse transliteration model to convert romanized Malayalam to native script using an encoder-decoder framework built with attention-based bidirectional Long Short Term Memory (Bi-LSTM)…

Computation and Language · Computer Science 2024-12-16 Bajiyo Baiju , Kavya Manohar , Leena G Pillai , Elizabeth Sherly

Digital text is increasing day by day on the internet. It is very challenging to classify a large and heterogeneous collection of data, which require improved information processing methods to organize text. To classify large size of…

Computation and Language · Computer Science 2021-07-08 Taimoor Ahmed Javed , Waseem Shahzad , Umair Arshad

Lemmatization is crucial for NLP tasks in morphologically rich languages with ambiguous orthography like Arabic, but existing tools face challenges due to inconsistent standards and limited genre coverage. This paper introduces two novel…

Computation and Language · Computer Science 2025-06-24 Mostafa Saeed , Nizar Habash

With the advent of digital optical scanners, a lot of paper-based books, textbooks, magazines, articles, and documents are being transformed into an electronic version that can be manipulated by a computer. For this purpose, OCR, short for…

Computation and Language · Computer Science 2012-04-03 Youssef Bassil , Mohammad Alwani

Arabic language and writing are now facing a resurgence of international normative solutions that challenge most of their local or network based operating principles. Even if the multilingual digital coding solutions, especially those…

Computers and Society · Computer Science 2017-03-14 Henri Hudrisier , Ben Henda Mokhtar

Morphological analysis is the study of the formation and structure of words. It plays a crucial role in various tasks in Natural Language Processing (NLP) and Computational Linguistics (CL) such as machine translation and text and speech…

Computation and Language · Computer Science 2020-05-22 Sina Ahmadi , Hossein Hassani

There are a lot of intensive researches on handwritten character recognition (HCR) for almost past four decades. The research has been done on some of popular scripts such as Roman, Arabic, Chinese and Indian. In this paper we present a…

Computer Vision and Pattern Recognition · Computer Science 2013-08-28 Aini Najwa Azmi , Dewi Nasien , Siti Mariyam Shamsuddin

Arabic morphological analysis is one of the essential stages in Arabic Natural Language Processing. In this paper we present an approach for Arabic morphological analysis. This approach is based on Arabic morphological automaton (AMAUT).…

Computation and Language · Computer Science 2011-01-31 Mourad Gridach , Noureddine Chenfour