English
Related papers

Related papers: Rule Based Stemmer in Urdu

200 papers

Representation of semantic information contained in the words is needed for any Arabic Text Mining applications. More precisely, the purpose is to better take into account the semantic dependencies between words expressed by the…

Computation and Language · Computer Science 2012-12-18 Hanane Froud , Abdelmonaim Lachkar , Said Alaoui Ouatik

Spelling errors are introduced in text either during typing, or when the user does not know the correct phoneme or grapheme. If a language contains complex words like sandhi where two or more morphemes join based on some rules, spell…

Computation and Language · Computer Science 2016-11-28 A N Akshatha , Chandana G Upadhyaya , Rajashekara S Murthy

We study the performance of Arabic text classification combining various techniques: (a) tfidf vs. dependency syntax, for feature selection and weighting; (b) class association rules vs. support vector machines, for classification. The…

Computation and Language · Computer Science 2014-10-21 Yannis Haralambous , Yassir Elidrissi , Philippe Lenca

This work presents a morphological analyzer for the Uzbek language using a finite state machine. The proposed methodology is a morphologic analysis of Uzbek words by using an affix striping to find a root and without including any lexicon.…

Computation and Language · Computer Science 2022-05-23 Maksud Sharipov , Ulugbek Salaev

Text detection in natural scene images has applications for autonomous driving, navigation help for elderly and blind people. However, the research on Urdu text detection is usually hindered by lack of data resources. We have developed a…

Computer Vision and Pattern Recognition · Computer Science 2022-09-29 Hazrat Ali

In this article, we present a rule-based approach for transliterating two mostly used orthographies in Sorani Kurdish. Our work consists of detecting a character in a word by removing the possible ambiguities and mapping it into the target…

Computation and Language · Computer Science 2018-11-27 Sina Ahmadi

A generate and test algorithm is described which parses a surface form into one or more lexical entries using linearly ordered phonological rules. This algorithm avoids the exponential expansion of search space which a naive parsing…

cmp-lg · Computer Science 2008-02-03 Michael Maxwell

In this article, an automated system is proposed for essay scoring in Arabic language for online exams based on stemming techniques and Levenshtein edit operations. An online exam has been developed on the proposed mechanisms, exploiting…

Information Retrieval · Computer Science 2016-11-10 Emad Fawzi Al-Shalabi

Differentiating intrinsic language words from transliterable words is a key step aiding text processing tasks involving different natural languages. We consider the problem of unsupervised separation of transliterable words from native…

Computation and Language · Computer Science 2018-03-28 Deepak P

Statistical error Correction technique is the most accurate and widely used approach today, but for a language like Sindhi which is a low resourced language the trained corpora's are not available, so the statistical techniques are not…

Computation and Language · Computer Science 2014-03-20 Zeeshan Bhatti , Imdad Ali Ismaili , Asad Ali Shaikh , Waseem Javaid

OCR algorithms have received a significant improvement in performance recently, mainly due to the increase in the capabilities of artificial intelligence algorithms. However, this advancement is not evenly distributed over all languages.…

Computer Vision and Pattern Recognition · Computer Science 2020-05-15 Atique Ur Rehman , Sibt Ul Hussain

Handwritten text recognition is an active research area in the field of deep learning and artificial intelligence to convert handwritten text into machine-understandable. A lot of work has been done for other languages, especially for…

Computer Vision and Pattern Recognition · Computer Science 2021-03-10 Muhammad Kashif

Neural Machine Translation models have replaced the conventional phrase based statistical translation methods since the former takes a generic, scalable, data-driven approach rather than relying on manual, hand-crafted features. The neural…

Computation and Language · Computer Science 2017-12-11 Mehreen Alam , Sibt ul Hussain

Idiomatic translation remains a significant challenge in machine translation, especially for low resource languages such as Urdu, and has received limited prior attention. To advance research in this area, we introduce the first evaluation…

Computation and Language · Computer Science 2025-10-21 Muhammad Farmal Khan , Mousumi Akter

This paper presents a method of stemming for the Arabian texts based on the linguistic techniques of the natural language processing. This method leans on the notion of scheme (one of the strong points of the morphology of the Arabian…

Computation and Language · Computer Science 2019-11-19 Sadik Bessou , Mohamed Louail , Allaoua Refoufi , Zehour Kadem , Mohamed Touahria

This study evaluates the feasibility of lightweight Whisper models (Tiny, Base, Small) for Urdu speech recognition in low-resource settings. Despite Urdu being the 10th most spoken language globally with over 230 million speakers, its…

Computation and Language · Computer Science 2025-08-14 Abdul Rehman Antall , Naveed Akhtar

English to Indian language machine translation poses the challenge of structural and morphological divergence. This paper describes English to Indian language statistical machine translation using pre-ordering and suffix separation. The…

Computation and Language · Computer Science 2018-08-03 Raj Nath Patel , Prakash B. Pimpale , M Sasikumar

Hindi being a highly inflectional language, FST (Finite State Transducer) based approach is most efficient for developing a morphological analyzer for this language. The work presented in this paper uses the SFST (Stuttgart Finite State…

Computation and Language · Computer Science 2012-07-24 Deepak Kumar , Manjeet Singh , Seema Shukla

Code-switching is a common phenomenon among people with diverse lingual background and is widely used on the internet for communication purposes. In this paper, we present a Recurrent Neural Network combined with the Attention Model for…

Computation and Language · Computer Science 2021-03-04 Aizaz Hussain , Muhammad Umair Arshad

Urdu, spoken by 230 million people worldwide, lacks dedicated transformer-based language models and curated corpora. While multilingual models provide limited Urdu support, they suffer from poor performance, high computational costs, and…

Computation and Language · Computer Science 2026-01-27 Syed Muhammad Ali , Hammad Sajid , Zainab Haider , Ali Muhammad Asad , Haya Fatima , Abdul Samad