中文
相关论文

相关论文: Accuracy of the Uzbek stop words detection: a case…

200 篇论文

The accurate syllabification of words plays a vital role in various Natural Language Processing applications. Syllabification is a versatile linguistic tool with applications in linguistic research, language technology, education, and…

计算与语言 · 计算机科学 2023-12-27 Ulugbek Salaev , Elmurod Kuriyozov , Gayrat Matlatipov

Automatic language identification is a natural language processing problem that tries to determine the natural language of a given content. In this paper we present a statistical method for automatic language identification of written text…

计算与语言 · 计算机科学 2018-06-15 Ciprian-Octavian Truică , Julien Velcin , Alexandru Boicea

The volume of information is increasing at an incredible rate with the rapid development of the Internet and electronic information services. Due to time constraints, we don't have the opportunity to read all this information. Even the task…

计算与语言 · 计算机科学 2023-03-02 Khabibulla Madatov , Shukurla Bekchanov , Jernej Vičič

This work presents a morphological analyzer for the Uzbek language using a finite state machine. The proposed methodology is a morphologic analysis of Uzbek words by using an affix striping to find a root and without including any lexicon.…

计算与语言 · 计算机科学 2022-05-23 Maksud Sharipov , Ulugbek Salaev

Lemmatization is one of the core concepts in natural language processing, thus creating a lemmatization tool is an important task. This paper discusses the construction of a lemmatization algorithm for the Uzbek language. The main purpose…

计算与语言 · 计算机科学 2022-10-31 Maksud Sharipov , Ogabek Sobirov

This research paper presents a part-of-speech (POS) annotated dataset and tagger tool for the low-resource Uzbek language. The dataset includes 12 tags, which were used to develop a rule-based POS-tagger tool. The corpus text used in the…

计算与语言 · 计算机科学 2023-03-02 Maksud Sharipov , Elmurod Kuriyozov , Ollabergan Yuldashev , Ogabek Sobirov

In this paper we present a rule-based stemming algorithm for the Uzbek language. Uzbek is an agglutinative language, so many words are formed by adding suffixes, and the number of suffixes is also large. For this reason, it is difficult to…

计算与语言 · 计算机科学 2022-10-31 Maksud Sharipov , Ollabergan Yuldashov

In this paper, we share the process of developing word embeddings for the Cyrillic variant of the Uzbek language. The result of our work is the first publicly available set of word vectors trained on the word2vec, GloVe, and fastText…

计算与语言 · 计算机科学 2020-10-01 B. Mansurov , A. Mansurov

Extracting useful information for sentiment analysis and classification problems from a big amount of user-generated feedback, such as restaurant reviews, is a crucial task of natural language processing, which is not only for customer…

计算与语言 · 计算机科学 2022-06-01 Sanatbek Matlatipov , Hulkar Rahimboeva , Jaloliddin Rajabov , Elmurod Kuriyozov

As Uzbek language is agglutinative, has many morphological features which words formed by combining root and affixes. Affixes play an important role in the morphological analysis of words, by adding additional meanings and grammatical…

计算与语言 · 计算机科学 2024-06-13 Ulugbek Salaev

One of the major challenges of an educational system is choosing appropriate content considering pupils' age and intellectual potential. In this article the experiment of primary school grades (from 1st to 4th grades) is considered for…

计算与语言 · 计算机科学 2023-03-21 Khabibulla Madatov , Sanatbek Matlatipov , Mersaid Aripov

Nowadays, creation of the tagged corpora is becoming one of the most important tasks of Natural Language Processing (NLP). There are not enough tagged corpora to build machine learning models for the low-resource Uzbek language. In this…

计算与语言 · 计算机科学 2022-10-28 Maksud Sharipov , Jamolbek Mattiev , Jasur Sobirov , Rustam Baltayev

Stop words, which are considered non-predictive, are often eliminated in natural language processing tasks. However, the definition of uninformative vocabulary is vague, so most algorithms use general knowledge-based stop lists to remove…

软件工程 · 计算机科学 2023-06-13 Yaohou Fan , Chetan Arora , Christoph Treude

Today text classification becomes critical task for concerned individuals for numerous purposes. Hence, several researches have been conducted to develop automatic text classification for national and international languages. However, the…

计算与语言 · 计算机科学 2023-05-09 Jawid Ahmad Baktash , Mursal Dawodi , Mohammad Zarif Joya , Nematullah Hassanzada

The effectiveness of three stop words lists for Arabic Information Retrieval---General Stoplist, Corpus-Based Stoplist, Combined Stoplist ---were investigated in this study. Three popular weighting schemes were examined: the inverse…

计算与语言 · 计算机科学 2017-02-08 Ibrahim Abu El-Khair

This paper investigates the role of text categorization in streamlining stopword extraction in natural language processing (NLP), specifically focusing on nine African languages alongside French. By leveraging the MasakhaNEWS, African…

Text classification systems will help to solve the text clustering problem in the Azerbaijani language. There are some text-classification applications for foreign languages, but we tried to build a newly developed system to solve this…

计算与语言 · 计算机科学 2020-01-01 Umid Suleymanov , Behnam Kiani Kalejahi , Elkhan Amrahov , Rashid Badirkhanli

Text classification is an important task in Natural Language Processing (NLP), where the goal is to categorize text data into predefined classes. In this study, we analyse the dataset creation steps and evaluation techniques of multi-label…

计算与语言 · 计算机科学 2023-03-01 Elmurod Kuriyozov , Ulugbek Salaev , Sanatbek Matlatipov , Gayrat Matlatipov

The study of natural language, especially Arabic, and mechanisms for the implementation of automatic processing is a fascinating field of study, with various potential applications. The importance of tools for natural language processing is…

计算与语言 · 计算机科学 2013-06-05 Riadh Bouslimi , Houda Amraoui

Stopwords are fundamental in Natural Language Processing (NLP) techniques for information retrieval. One of the common tasks in preprocessing of text data is the removal of stopwords. Currently, while high-resource languages like English…

‹ 上一页 1 2 3 10 下一页 ›