English
Related papers

Related papers: Unicode Normalization and Grapheme Parsing of Indi…

200 papers

Arabic language and writing are now facing a resurgence of international normative solutions that challenge most of their local or network based operating principles. Even if the multilingual digital coding solutions, especially those…

Computers and Society · Computer Science 2017-03-14 Henri Hudrisier , Ben Henda Mokhtar

This paper presents a novel approach towards Indic handwritten word recognition using zone-wise information. Because of complex nature due to compound characters, modifiers, overlapping and touching, etc., character segmentation and…

Computer Vision and Pattern Recognition · Computer Science 2017-08-02 Partha Pratim Roy , Ayan Kumar Bhunia , Ayan Das , Prasenjit Dey , Umapada Pal

We develop a general finite-alphabet framework for Euler-type sums based on the notion of a monoidal alphabet. An alphabet of summand letters is called monoidal when it is closed under pointwise multiplication, thereby inducing the usual…

General Mathematics · Mathematics 2026-05-22 Jayanta Phadikar

This paper investigates the employment of various encoders in text transformation, converting characters into bytes. It discusses local encoders such as ASCII and GB-2312, which encode specific characters into shorter bytes, and universal…

Computation and Language · Computer Science 2023-07-12 Changshang Xue

Leveraging Graph Neural Networks (GNNs) as graph encoders and aligning the resulting representations with Large Language Models (LLMs) through alignment instruction tuning has become a mainstream paradigm for constructing Graph Language…

Machine Learning · Computer Science 2026-05-13 Haibo Chen , Xin Wang , Jiaheng Chao , Ling Feng , Wenwu Zhu

The cursive nature of multilingual characters segmentation and recognition of Arabic, Persian, Urdu languages have attracted researchers from academia and industry. However, despite several decades of research, still multilingual characters…

Computer Vision and Pattern Recognition · Computer Science 2019-04-19 Amjad Rehman , Majid Harouni , Tanzila Saba

Universal Dependencies (UD) offer a uniform cross-lingual syntactic representation, with the aim of advancing multilingual applications. Recent work shows that semantic parsing can be accomplished by transforming syntactic dependencies to…

Computation and Language · Computer Science 2017-08-30 Siva Reddy , Oscar Täckström , Slav Petrov , Mark Steedman , Mirella Lapata

Transliteration, the process of mapping text from one script to another, plays a crucial role in multilingual natural language processing, especially within linguistically diverse contexts such as India. Despite significant advancements…

Computation and Language · Computer Science 2025-05-27 Gulfarogh Azam , Mohd Sadique , Saif Ali , Mohammad Nadeem , Erik Cambria , Shahab Saquib Sohail , Mohammad Sultan Alam

State-of-the-art speech recognition systems rely heavily on three basic components: an acoustic model, a pronunciation lexicon and a language model. To build these components, a researcher needs linguistic as well as technical expertise,…

Computation and Language · Computer Science 2018-03-06 Haris Bin Zia , Agha Ali Raza , Awais Athar

This review paper provides a comprehensive overview of large language model (LLM) research directions within Indic languages. Indic languages are those spoken in the Indian subcontinent, including India, Pakistan, Bangladesh, Sri Lanka,…

Computation and Language · Computer Science 2024-06-17 Sankalp KJ , Vinija Jain , Sreyoshi Bhaduri , Tamoghna Roy , Aman Chadha

Text normalization is a ubiquitous process that appears as the first step of many Natural Language Processing problems. However, previous Deep Learning approaches have suffered from so-called silly errors, which are undetectable on…

Computation and Language · Computer Science 2019-03-08 Adrián Javaloy Bornás , Ginés García Mateos

The work presented here involves the design of a Multi Layer Perceptron (MLP) based classifier for recognition of handwritten Bangla alphabet using a 76 element feature set Bangla is the second most popular script and language in the Indian…

Computer Vision and Pattern Recognition · Computer Science 2012-03-06 Subhadip Basu , Nibaran Das , Ram Sarkar , Mahantapas Kundu , Mita Nasipuri , Dipak Kumar Basu

Transliteration is a task in the domain of NLP where the output word is a similar-sounding word written using the letters of any foreign language. Today this system has been developed for several language pairs that involve English as…

Computation and Language · Computer Science 2022-08-24 Yash Raj , Bhavesh Laddagiri

Urdu is a challenging language because of, first, its Perso-Arabic script and second, its morphological system having inherent grammatical forms and vocabulary of Arabic, Persian and the native languages of South Asia. This paper describes…

Computation and Language · Computer Science 2022-04-08 Muhammad Humayoun , Harald Hammarström , Aarne Ranta

Tokenisation is the first step in almost all NLP tasks, and state-of-the-art transformer-based language models all use subword tokenisation algorithms to process input text. Existing algorithms have problems, often producing tokenisations…

Computation and Language · Computer Science 2022-10-25 Edward Gow-Smith , Harish Tayyar Madabushi , Carolina Scarton , Aline Villavicencio

This paper presents a novel approach to generate synthetic dataset for handwritten word recognition systems. It is difficult to recognize handwritten scripts for which sufficient training data is not readily available or it may be expensive…

Computer Vision and Pattern Recognition · Computer Science 2018-04-18 Partha Pratim Roy , Akash Mohta , Bidyut B. Chaudhuri

Social media user-generated text is actually the main resource for many NLP tasks. This text however, does not follow the standard rules of writing. Moreover, the use of dialect such as Moroccan Arabic in written communications increases…

Computation and Language · Computer Science 2022-06-22 Randa Zarnoufi , Walid Bachri , Hamid Jaafar , Mounia Abik

Standard natural language processing (NLP) pipelines operate on symbolic representations of language, which typically consist of sequences of discrete tokens. However, creating an analogous representation for ancient logographic writing…

Computation and Language · Computer Science 2026-01-29 Danlu Chen , Freda Shi , Aditi Agarwal , Jacobo Myerston , Taylor Berg-Kirkpatrick

The performance of Language Models (LMs) on low-resource, morphologically rich languages like Sinhala remains largely unexplored, particularly regarding script variation in digital communication. Sinhala exhibits script duality, with…

Computation and Language · Computer Science 2026-05-11 Minuri Rajapakse , Ruvan Weerasinghe

Comprehensively searching for words in Sanskrit E-text is a non-trivial problem because words could change their forms in different contexts. One such context is sandhi or euphonic conjunctions, which cause a word to change owing to the…

Computation and Language · Computer Science 2019-08-17 S. V. Kasmir Raja , V. Rajitha , Meenakshi Lakshmanan
‹ Prev 1 3 4 5 6 7 10 Next ›