English
Related papers

Related papers: Mavericks at NADI 2023 Shared Task: Unravelling Re…

200 papers

Dialects introduce syntactic and lexical variations in language that occur in regional or social groups. Most NLP methods are not sensitive to such variations. This may lead to unfair behavior of the methods, conveying negative bias towards…

Computation and Language · Computer Science 2024-06-17 Maximilian Spliethöver , Sai Nikhil Menon , Henning Wachsmuth

Text classification has been one of the earliest problems in NLP. Over time the scope of application areas has broadened and the difficulty of dealing with new areas (e.g., noisy social media content) has increased. The problem-solving…

Computation and Language · Computer Science 2020-11-10 Tanvirul Alam , Akib Khan , Firoj Alam

Native language identification (NLI) is the task of automatically identifying the native language (L1) of an individual based on their language production in a learned language. It is useful for a variety of purposes including marketing,…

Computation and Language · Computer Science 2022-11-21 Ahmet Yavuz Uluslu , Gerold Schneider

Since their inception, transformer-based language models have led to impressive performance gains across multiple natural language processing tasks. For Arabic, the current state-of-the-art results on most datasets are achieved by the…

Computation and Language · Computer Science 2021-03-11 Amey Hengle , Atharva Kshirsagar , Shaily Desai , Manisha Marathe

Morphological tagging is challenging for morphologically rich languages due to the large target space and the need for more training data to minimize model sparsity. Dialectal variants of morphologically rich languages suffer more as they…

Computation and Language · Computer Science 2019-10-29 Nasser Zalmout , Nizar Habash

This paper presents an ensemble system combining the output of multiple SVM classifiers to native language identification (NLI). The system was submitted to the NLI Shared Task 2017 fusion track which featured students essays and spoken…

Computation and Language · Computer Science 2017-07-25 Marcos Zampieri , Alina Maria Ciobanu , Liviu P. Dinu

This paper presents a novel method that allows a machine learning algorithm following the transformation-based learning paradigm \cite{brill95:tagging} to be applied to multiple classification tasks by training jointly and simultaneously on…

Computation and Language · Computer Science 2007-05-23 Radu Florian , Grace Ngai

We propose a general method to break down a main complex task into a set of intermediary easier sub-tasks, which are formulated in natural language as binary questions related to the final target task. Our method allows for representing…

Computation and Language · Computer Science 2024-02-02 Felipe Urrutia , Cristian Buc , Valentin Barriere

Dialect Identification is a crucial task for localizing various Large Language Models. This paper outlines our approach to the VarDial 2023 shared task. Here we have to identify three or two dialects from three languages each which results…

Computation and Language · Computer Science 2023-03-29 Ankit Vaidya , Aditya Kane

This paper presents our system developed for the SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages. The shared task aims at measuring the semantic textual relatedness between pairs of sentences, with a focus…

Computation and Language · Computer Science 2024-06-10 Miaoran Zhang , Mingyang Wang , Jesujoba O. Alabi , Dietrich Klakow

We present MSAs winning system for the BAREC 2025 Shared Task on fine-grained Arabic readability assessment, achieving first place in six of six tracks. Our approach is a confidence-weighted ensemble of four complementary transformer models…

Computation and Language · Computer Science 2025-09-15 Mohamed Basem , Mohamed Younes , Seif Ahmed , Abdelrahman Moustafa

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects alongside Modern…

Computation and Language · Computer Science 2024-12-19 Basel Mousi , Nadir Durrani , Fatema Ahmad , Md. Arid Hasan , Maram Hasanain , Tameem Kabbani , Fahim Dalvi , Shammur Absar Chowdhury , Firoj Alam

Named Entity Recognition (NER) is a task in Natural Language Processing (NLP) that aims to identify and classify entities in text into predefined categories. However, when applied to Arabic data, NER encounters unique challenges stemming…

Computation and Language · Computer Science 2024-08-08 Ahmed Abdou , Tasneem Mohsen

The acoustic and linguistic features are important cues for the spoken language identification (LID) task. Recent advanced LID systems mainly use acoustic features that lack the usage of explicit linguistic feature encoding. In this paper,…

Computation and Language · Computer Science 2022-08-01 Peng Shen , Xugang Lu , Hisashi Kawai

This paper describes the approach of the UniBuc - NLP team in tackling the SemEval 2024 Task 8: Multigenerator, Multidomain, and Multilingual Black-Box Machine-Generated Text Detection. We explored transformer-based and hybrid deep learning…

Computation and Language · Computer Science 2024-05-29 Teodor-George Marchitan , Claudiu Creanga , Liviu P. Dinu

The prevalence of toxic content on social media platforms, such as hate speech, offensive language, and misogyny, presents serious challenges to our interconnected society. These challenging issues have attracted widespread attention in…

Computation and Language · Computer Science 2022-06-20 Abdelkader El Mahdaouy , Abdellah El Mekki , Ahmed Oumar , Hajar Mousannif , Ismail Berrada

Speech acts are a speakers actions when performing an utterance within a conversation, such as asking, recommending, greeting, or thanking someone, expressing a thought, or making a suggestion. Understanding speech acts helps interpret the…

Computation and Language · Computer Science 2024-02-01 Khadejaa Alshehri , Areej Alhothali , Nahed Alowidi

We present ADI-20, an extension of the previously published ADI-17 Arabic Dialect Identification (ADI) dataset. ADI-20 covers all Arabic-speaking countries' dialects. It comprises 3,556 hours from 19 Arabic dialects in addition to Modern…

Computation and Language · Computer Science 2025-11-14 Haroun Elleuch , Salima Mdhaffar , Yannick Estève , Fethi Bougares

Current Machine Translation (MT) systems for Arabic often struggle to account for dialectal diversity, frequently homogenizing dialectal inputs into Modern Standard Arabic (MSA) and offering limited user control over the target vernacular.…

Computation and Language · Computer Science 2026-04-09 Afroza Nowshin , Prithweeraj Acharjee Porag , Haziq Jeelani , Fayeq Jeelani Syed

This paper presents six document classification models using the latest transformer encoders and a high-performing ensemble model for a task of offensive language identification in social media. For the individual models, deep transformer…

Computation and Language · Computer Science 2020-07-22 Xiangjue Dong , Jinho D. Choi