中文
相关论文

相关论文: Rule Based Stemmer in Urdu

200 篇论文

Fully data-driven, deep learning-based models are usually designed as language-independent and have been shown to be successful for many natural language processing tasks. However, when the studied language is low-resourced and the amount…

计算与语言 · 计算机科学 2022-09-21 Şaziye Betül Özateş , Arzucan Özgür , Tunga Güngör , Balkız Öztürk

Text normalization is an essential preprocessing step in many natural language processing (NLP) tasks, and stemming is one such normalization technique that reduces words to their base or root form. However, evaluating stemming methods is…

计算与语言 · 计算机科学 2025-11-26 Md Abdullah Al Kafi , Raka Moni , Sumit Kumar Banshal

The conventional natural language processing approaches are not accustomed to the social media text due to colloquial discourse and non-homogeneous characteristics. Significantly, the language identification in a multilingual document is…

计算与语言 · 计算机科学 2021-06-30 M Zeeshan Ansari , Tanvir Ahmad , M M Sufyan Beg , Asma Ikram

We present a language independent, unsupervised approach for transforming word embeddings from source language to target language using a transformation matrix. Our model handles the problem of data scarcity which is faced by many languages…

Social media has become an essential part of the digital age, serving as a platform for communication, interaction, and information sharing. Celebrities are among the most active users and often reveal aspects of their personal and…

社会与信息网络 · 计算机科学 2025-10-15 Muhammad Hamza , Rizwan Jafar

In this paper we present our work on a case study between Statistical Machien Transaltion (SMT) and Rule-Based Machine Translation (RBMT) systems on English-Indian langugae and Indian to Indian langugae perspective. Main objective of our…

计算与语言 · 计算机科学 2017-08-16 Sreelekha S

Online forums or message boards are rich knowledge-based communities. In these communities, thread retrieval is an essential tool facilitating information access. However, the issue on thread search is how to combine evidence from text…

信息检索 · 计算机科学 2013-01-17 Ameer Tawfik Albaham , Naomie Salim

Based on an annotated multimedia corpus, television series Mar{\=a}y{\=a} 2013, we dig into the question of ''automatic standardization'' of Arabic dialects for machine translation. Here we distinguish between rule-based machine translation…

计算与语言 · 计算机科学 2023-01-10 Abidrabbo Alnassan

Document level Urdu Sentiment Analysis (SA) is a challenging Natural Language Processing (NLP) task as it deals with large documents in a resource-poor language. In large documents, there are ample amounts of words that exhibit different…

计算与语言 · 计算机科学 2025-01-30 Ammarah Irum , M. Ali Tahir

Reduplication and repetition, though similar in form, serve distinct linguistic purposes. Reduplication is a deliberate morphological process used to express grammatical, semantic, or pragmatic nuances, while repetition is often…

计算与语言 · 计算机科学 2024-07-12 Arif Ahmad , Mothika Gayathri Khyathi , Pushpak Bhattacharyya

As Uzbek language is agglutinative, has many morphological features which words formed by combining root and affixes. Affixes play an important role in the morphological analysis of words, by adding additional meanings and grammatical…

计算与语言 · 计算机科学 2024-06-13 Ulugbek Salaev

Fast and accurate spoken content retrieval is vital for applications such as voice search. Query-by-Example Spoken Term Detection (STD) involves retrieving matching segments from an audio database given a spoken query. Token-based STD…

音频与语音处理 · 电气工程与系统科学 2026-02-19 Anup Singh , Vipul Arora , Kris Demuynck

Phrase-based Statistical models are more commonly used as they perform optimally in terms of both, translation quality and complexity of the system. Hindi and in general all Indian languages are morphologically richer than English. Hence,…

计算与语言 · 计算机科学 2017-09-19 Sreelekha S , Pushpak Bhattacharyya

Information Retrieval (IR) allows the storage, management, processing and retrieval of information, documents, websites, etc. Building an IR system for any language is imperative. This is evident through the massive conducted efforts to…

信息检索 · 计算机科学 2018-01-16 Bilal Abu-Salih

In order to accelerate the performance of various Natural Language Processing tasks for Roman Urdu, this paper for the very first time provides 3 neural word embeddings prepared using most widely used approaches namely Word2vec, FastText,…

计算与语言 · 计算机科学 2020-03-13 Faiza Memood , Muhammad Usman Ghani , Muhammad Ali Ibrahim , Rehab Shehzadi , Muhammad Nabeel Asim

In Automatic Text Summarization, preprocessing is an important phase to reduce the space of textual representation. Classically, stemming and lemmatization have been widely used for normalizing words. However, even using normalization on…

信息检索 · 计算机科学 2012-09-17 Juan-Manuel Torres-Moreno

In this paper we explore the problem of document summarization in Persian language from two distinct angles. In our first approach, we modify a popular and widely cited Persian document summarization framework to see how it works on a…

计算与语言 · 计算机科学 2016-06-13 Saeid Parvandeh , Shibamouli Lahiri , Fahimeh Boroumand

Text detection in natural scene images for content analysis is an interesting task. The research community has seen some great developments for English/Mandarin text detection. However, Urdu text extraction in natural scene images is a task…

计算机视觉与模式识别 · 计算机科学 2021-09-17 Hazrat Ali , Khalid Iqbal , Ghulam Mujtaba , Ahmad Fayyaz , Mohammad Farhad Bulbul , Fazal Wahab Karam , Ali Zahir

Evaluation plays a crucial role in development of Machine translation systems. In order to judge the quality of an existing MT system i.e. if the translated output is of human translation quality or not, various automatic metrics exist. We…

计算与语言 · 计算机科学 2014-04-08 Aditi Kalyani , Hemant Kumud , Shashi Pal Singh , Ajai Kumar , Hemant Darbari

Word segmentation is a low-level NLP task that is non-trivial for a considerable number of languages. In this paper, we present a sequence tagging framework and apply it to word segmentation for a wide range of languages with different…

计算与语言 · 计算机科学 2018-07-10 Yan Shao , Christian Hardmeier , Joakim Nivre