English
Related papers

Related papers: An Amharic News Text classification Dataset

200 papers

Over the last few years, Text classification is one of the fundamental tasks in natural language processing (NLP) in which the objective is to categorize text documents into one of the predefined classes. The news is full of our life.…

Computation and Language · Computer Science 2022-01-26 Ke Yahan , Ruyi Qu , Lu Xiaoxia

Knowledge is central to human and scientific developments. Natural Language Processing (NLP) allows automated analysis and creation of knowledge. Data is a crucial NLP and machine learning ingredient. The scarcity of open datasets is a…

Computation and Language · Computer Science 2022-10-19 Istiak Ahmad , Fahad AlQurashi , Rashid Mehmood

We introduce ALHD, the first large-scale comprehensive Arabic dataset explicitly designed to distinguish between human- and LLM-generated texts. ALHD spans three genres (news, social media, reviews), covering both MSA and dialectal Arabic,…

Computation and Language · Computer Science 2025-10-23 Ali Khairallah , Arkaitz Zubiaga

The NLP pipeline has evolved dramatically in the last few years. The first step in the pipeline is to find suitable annotated datasets to evaluate the tasks we are trying to solve. Unfortunately, most of the published datasets lack metadata…

Computation and Language · Computer Science 2021-10-14 Zaid Alyafeai , Maraim Masoud , Mustafa Ghaleb , Maged S. Al-shaibani

While natural language processing tools have been developed extensively for some of the world's languages, a significant portion of the world's over 7000 languages are still neglected. One reason for this is that evaluation datasets do not…

Computation and Language · Computer Science 2024-06-05 Chunlan Ma , Ayyoob ImaniGooghari , Haotian Ye , Renhao Pei , Ehsaneddin Asgari , Hinrich Schütze

For high-resource languages like English, text classification is a well-studied task. The performance of modern NLP models easily achieves an accuracy of more than 90% in many standard datasets for text classification in English (Xie et…

Computation and Language · Computer Science 2022-06-06 Dawei Zhu , Michael A. Hedderich , Fangzhou Zhai , David Ifeoluwa Adelani , Dietrich Klakow

This paper reports some difficulties and some results when using dense retrievers on Amharic, one of the low-resource languages spoken by 120 millions populations. The efforts put and difficulties faced by University Addis Ababa toward…

Information Retrieval · Computer Science 2025-03-25 Tilahun Yeshambel , Moncef Garouani , Serge Molina , Josiane Mothe

This paper addresses the classification of Arabic text data in the field of Natural Language Processing (NLP), with a particular focus on Natural Language Inference (NLI) and Contradiction Detection (CD). Arabic is considered a…

Computation and Language · Computer Science 2023-07-28 Mohammad Majd Saad Al Deen , Maren Pielka , Jörn Hees , Bouthaina Soulef Abdou , Rafet Sifa

Recent advances in word embeddings and language models use large-scale, unlabelled data and self-supervised learning to boost NLP performance. Multilingual models, often trained on web-sourced data like Wikipedia, face challenges: few…

Computation and Language · Computer Science 2025-07-02 David Ifeoluwa Adelani

The widespread absence of diacritical marks in Arabic text poses a significant challenge for Arabic natural language processing (NLP). This paper explores instances of naturally occurring diacritics, referred to as "diacritics in the wild,"…

Computation and Language · Computer Science 2024-06-11 Salman Elgamal , Ossama Obeid , Tameem Kabbani , Go Inoue , Nizar Habash

Recently, pre-trained transformer-based architectures have proven to be very efficient at language modeling and understanding, given that they are trained on a large enough corpus. Applications in language generation for Arabic are still…

Computation and Language · Computer Science 2021-03-09 Wissam Antoun , Fady Baly , Hazem Hajj

Few-shot learning is an important, but challenging problem of machine learning aimed at learning from only fewer labeled training examples. It has become an active area of research due to deep learning requiring huge amounts of labeled…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Mesay Samuel , Lars Schmidt-Thieme , DP Sharma , Abiot Sinamo , Abey Bruck

The growing importance of culturally-aware natural language processing systems has led to an increasing demand for resources that capture sociopragmatic phenomena across diverse languages. Nevertheless, Arabic-language resources for…

Text classification has become a crucial task in various fields, leading to a significant amount of research on developing automated text classification systems for national and international languages. However, there is a growing need for…

Computation and Language · Computer Science 2023-05-08 Mursal Dawodi , Jawid Ahmad Baktash

Training data for machine learning models can come from many different sources, which can be of dubious quality. For resource-rich languages like English, there is a lot of data available, so we can afford to throw out the dubious data. For…

Computation and Language · Computer Science 2021-03-31 Andrew Zupon , Evan Crew , Sandy Ritchie

Recent progress in text classification has been focused on high-resource languages such as English and Chinese. For low-resource languages, amongst them most African languages, the lack of well-annotated data and effective preprocessing, is…

Computation and Language · Computer Science 2020-10-26 Rubungo Andre Niyongabo , Hong Qu , Julia Kreutzer , Li Huang

Research in Natural Language Processing (NLP) has increasingly become important due to applications such as text classification, text mining, sentiment analysis, POS tagging, named entity recognition, textual entailment, and many others.…

Artificial Intelligence · Computer Science 2022-10-21 Istiak Ahmad , Fahad AlQurashi , Rashid Mehmood

This study examines the digital representation of African languages and the challenges this presents for current language detection tools. We evaluate their performance on Yoruba, Kinyarwanda, and Amharic. While these languages are spoken…

Computation and Language · Computer Science 2026-01-27 Edward Ajayi , Eudoxie Umwari , Mawuli Deku , Prosper Singadi , Jules Udahemuka , Bekalu Tadele , Chukuemeka Edeh

Machine translation (MT) systems are now able to provide very accurate results for high resource language pairs. However, for many low resource languages, MT is still under active research. In this paper, we develop and share a dataset to…

Computation and Language · Computer Science 2020-04-01 Asmelash Teka Hadgu , Adam Beaudoin , Abel Aregawi

Fake news detection research is still in the early stage as this is a relatively new phenomenon in the interest raised by society. Machine learning helps to solve complex problems and to build AI systems nowadays and especially in those…

Computation and Language · Computer Science 2022-01-20 Sajjad Ahmed , Knut Hinkelmann , Flavio Corradini