English
Related papers

Related papers: A Tale of Two Scripts: Transliteration and Post-Co…

200 papers

Automatic Arabic Dialect Identification (ADI) of text has gained great popularity since it was introduced in the early 2010s. Multiple datasets were developed, and yearly shared tasks have been running since 2018. However, ADI systems are…

Computation and Language · Computer Science 2023-10-23 Amr Keleg , Walid Magdy

We present AraLingBench: a fully human annotated benchmark for evaluating the Arabic linguistic competence of large language models (LLMs). The benchmark spans five core categories: grammar, morphology, spelling, reading comprehension, and…

Computation and Language · Computer Science 2025-12-10 Mohammad Zbeeb , Hasan Abed Al Kader Hammoud , Sina Mukalled , Nadine Rizk , Fatima Karnib , Issam Lakkis , Ammar Mohanna , Bernard Ghanem

This paper introduces a pioneering English-Azerbaijani (Arabic Script) parallel corpus, designed to bridge the technological gap in language learning and machine translation (MT) for under-resourced languages. Consisting of 548,000 parallel…

The effectiveness of the BERT model on multiple linguistic tasks has been well documented. On the other hand, its potentials for narrow and specific domains such as Legal, have not been fully explored. In this paper, we examine how BERT can…

Computation and Language · Computer Science 2022-10-18 Muhammad AL-Qurishi , Sarah AlQaseemi , Riad Soussi

The paper proposes an algorithm for the script recognition based on the texture characteristics. The image texture is achieved by coding each letter with the equivalent script type (number code) according to its position in the text line.…

Computer Vision and Pattern Recognition · Computer Science 2015-09-01 Darko Brodic , Zoran N. Milivojevic , Alessia Amelio

A hybrid approach for the transliteration of Algerian Arabizi: A primary study In this paper, we present a hybrid approach for the transliteration of the Algerian Arabizi. We define a set of rules enable us the passage from Arabizi to…

Computation and Language · Computer Science 2018-08-13 Imane Guellil , Faical Azouaou , Fodil Benali , Ala-Eddine Hachani , Houda Saadane

Question semantic similarity is a challenging and active research problem that is very useful in many NLP applications, such as detecting duplicate questions in community question answering platforms such as Quora. Arabic is considered to…

Computation and Language · Computer Science 2019-09-23 Hesham Al-Bataineh , Wael Farhan , Ahmad Mustafa , Haitham Seelawi , Hussein T. Al-Natsheh

This paper explores the use of a learned classifier for post-OCR text correction. Experiments with the Arabic language show that this approach, which integrates a weighted confusion matrix and a shallow language model, improves the vast…

Information Retrieval · Computer Science 2020-06-11 Ido Kissos , Nachum Dershowitz

Research into statistical parsing for English has enjoyed over a decade of successful results. However, adapting these models to other languages has met with difficulties. Previous comparative work has shown that Modern Arabic is one of the…

Computation and Language · Computer Science 2015-10-27 Kais Dukes

While Knowledge Editing (KE) has been widely explored in English, its behavior in morphologically rich languages like Arabic remains underexamined. In this work, we present the first study of Arabic KE. We evaluate four methods (ROME,…

Computation and Language · Computer Science 2025-11-04 Basel Mousi , Nadir Durrani , Fahim Dalvi

With the development of electronic media and the heterogeneity of Arabic data on the Web, the idea of building a clean corpus for certain applications of natural language processing, including machine translation, information retrieval,…

Computation and Language · Computer Science 2017-09-28 Wided Bakari , Patrice Bellot , Mahmoud Neji

The national languages of Senegal, like those of West Africa country in general, are written with two alphabets : the Latin alphabet that draws its strength from official decreesm and the completed Arabic script (Ajami), widespread and well…

Computation and Language · Computer Science 2020-05-07 El hadji M. Fall , El hadji M. Nguer , Bao Diop Sokhna , Mouhamadou Khoule , Mathieu Mangeot , Mame T. Cisse

This paper presents a novel Dialectal Sound and Vowelization Recovery framework, designed to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages, that extends beyond its standard orthographic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-06 Yassine El Kheir , Hamdy Mubarak , Ahmed Ali , Shammur Absar Chowdhury

With the expanding growth of Arabic electronic data on the web, extracting information, which is actually one of the major challenges of the question-answering, is essentially used for building corpus of documents. In fact, building a…

Information Retrieval · Computer Science 2018-05-24 Patrice Bellot , Wided Bakari , Mahmoud Neji

In Natural Language Processing (NLP), one of the most important tasks is text-to-SQL semantic parsing, which focuses on enabling users to interact with the database in a more natural manner. In recent years, text-to-SQL has made significant…

Computation and Language · Computer Science 2024-02-26 Saleh Almohaimeed , Saad Almohaimeed , Mansour Al Ghanim , Liqiang Wang

We demonstrate that it is feasible to diacritize Hebrew script without any human-curated resources other than plain diacritized text. We present NAKDIMON, a two-layer character level LSTM, that performs on par with much more complicated…

Computation and Language · Computer Science 2022-05-11 Elazar Gershuni , Yuval Pinter

Many automatic translation works have been addressed between major European language pairs, by taking advantage of large scale parallel corpora, but very few research works are conducted on the Amharic-Arabic language pair due to its…

Computation and Language · Computer Science 2020-01-01 Ibrahim Gashaw , H L Shashirekha

With the growing number of textual resources available, the ability to understand them becomes critical. An essential first step in understanding these sources is the ability to identify the part of speech in each sentence. Arabic is a…

Computation and Language · Computer Science 2013-07-15 Bilel Ben Ali , Fethi Jarray

Yor\`ub\'a is a widely spoken West African language with a writing system rich in tonal and orthographic diacritics. With very few exceptions, diacritics are omitted from electronic texts, due to limited device and application support.…

Computation and Language · Computer Science 2018-10-31 Iroro Orife

Recognition of Arabic-like scripts such as Persian and Urdu is more challenging than Latin-based scripts. This is due to the presence of a two-dimensional structure, context-dependent character shapes, spaces and overlaps, and placement of…

Computer Vision and Pattern Recognition · Computer Science 2021-11-09 Tayaba Anjum , Nazar Khan