English
Related papers

Related papers: Classifier Ensembles for Dialect and Language Vari…

200 papers

Most languages of the world pose low-resource challenges to natural language processing models. With multilingual training, knowledge can be shared among languages. However, not all languages positively influence each other and it is an…

Machine Learning · Computer Science 2023-10-25 Mingyang Wang , Heike Adel , Lukas Lange , Jannik Strötgen , Hinrich Schütze

Language identification is an important first step in many IR and NLP applications. Most publicly available language identification datasets, however, are compiled under the assumption that the gold label of each instance is determined by…

Computation and Language · Computer Science 2023-03-03 Marcos Zampieri , Kai North , Tommi Jauhiainen , Mariano Felice , Neha Kumari , Nishant Nair , Yash Bangera

Due to the scarcity of labeled dialectal speech, audio dialect classification is a challenging task for most languages, including Swiss German. In this work, we explore the ability of large language models (LLMs) as agents in understanding…

Computation and Language · Computer Science 2026-04-01 Tobias Bystrich , Lukas Hamm , Maria Hassan , Lea Fischbach , Lucie Flek , Akbar Karimi

Dialogue systems for Automatic Differential Diagnosis (ADD) have a wide range of real-life applications. These dialogue systems are promising for providing easy access and reducing medical costs. Building end-to-end ADD dialogue systems…

Computation and Language · Computer Science 2023-08-17 Srija Macherla , Man Luo , Mihir Parmar , Chitta Baral

This study aims at investigating the effect of applying single learner machine learning approach and ensemble machine learning approach for offensive language detection on Arabic language. Classifying Arabic social media text is a very…

Computation and Language · Computer Science 2020-05-20 Fatemah Husain

This paper presents an overview of the Arabic Natural Language Understanding (ArabicNLU 2024) shared task, focusing on two subtasks: Word Sense Disambiguation (WSD) and Location Mention Disambiguation (LMD). The task aimed to evaluate the…

Computation and Language · Computer Science 2024-07-31 Mohammed Khalilia , Sanad Malaysha , Reem Suwaileh , Mustafa Jarrar , Alaa Aljabari , Tamer Elsayed , Imed Zitouni

The use of multilingual language models for tasks in low and high-resource languages has been a success story in deep learning. In recent times, Arabic has been receiving widespread attention on account of its dialectal variance. While…

Computation and Language · Computer Science 2022-11-09 Soumajyoti Sarkar , Kaixiang Lin , Sailik Sengupta , Leonard Lausen , Sheng Zha , Saab Mansour

Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we propose the LinTO…

Computation and Language · Computer Science 2025-04-04 Hedi Naouara , Jean-Pierre Lorré , Jérôme Louradour

Current Machine Translation (MT) systems for Arabic often struggle to account for dialectal diversity, frequently homogenizing dialectal inputs into Modern Standard Arabic (MSA) and offering limited user control over the target vernacular.…

Computation and Language · Computer Science 2026-04-09 Afroza Nowshin , Prithweeraj Acharjee Porag , Haziq Jeelani , Fayeq Jeelani Syed

The diverse nature of dialects presents challenges for models trained on specific linguistic patterns, rendering them susceptible to errors when confronted with unseen or out-of-distribution (OOD) data. This study introduces a novel…

Computation and Language · Computer Science 2024-06-27 Yaqian Hao , Chenguang Hu , Yingying Gao , Shilei Zhang , Junlan Feng

Large language models (LLMs) have shown remarkable performance across various sentence-based linguistic phenomena, yet their ability to capture cross-sentence paradigmatic patterns, such as verb alternations, remains underexplored. In this…

Computation and Language · Computer Science 2026-03-17 Giuseppe Samo , Paola Merlo

Since their inception, transformer-based language models have led to impressive performance gains across multiple natural language processing tasks. For Arabic, the current state-of-the-art results on most datasets are achieved by the…

Computation and Language · Computer Science 2021-03-11 Amey Hengle , Atharva Kshirsagar , Shaily Desai , Manisha Marathe

Recently, string kernels have obtained state-of-the-art results in various text classification tasks such as Arabic dialect identification or native language identification. In this paper, we apply two simple yet effective transductive…

Computation and Language · Computer Science 2018-09-03 Radu Tudor Ionescu , Andrei M. Butnaru

Dialectal Arabic (DA) varieties are under-served by language technologies, particularly large language models (LLMs). This trend threatens to exacerbate existing social inequalities and limits LLM applications, yet the research community…

Computation and Language · Computer Science 2025-01-07 Nathaniel R. Robinson , Shahd Abdelmoneim , Kelly Marchisio , Sebastian Ruder

This paper addresses critical gaps in Arabic language model evaluation by establishing comprehensive theoretical guidelines and introducing a novel evaluation framework. We first analyze existing Arabic evaluation datasets, identifying…

Computation and Language · Computer Science 2025-06-03 Serry Sibaee , Omer Nacar , Adel Ammar , Yasser Al-Habashi , Abdulrahman Al-Batati , Wadii Boulila

Large Language Models (LLMs) have shown remarkable capabilities, not only in generating human-like text, but also in acquiring knowledge. This highlights the need to go beyond the typical Natural Language Processing downstream benchmarks…

Computation and Language · Computer Science 2025-01-03 Ahmad Mustapha , Hadi Al-Khansa , Hadi Al-Mubasher , Aya Mourad , Ranam Hamoud , Hasan El-Husseini , Marwah Al-Sakkaf , Mariette Awad

Language Identification (LID) is a core task in multilingual NLP, yet current systems often overfit to clean, monolingual data. This work introduces DIVERS-BENCH, a comprehensive evaluation of state-of-the-art LID models across diverse…

Computation and Language · Computer Science 2025-09-23 Jessica Ojo , Zina Kamel , David Ifeoluwa Adelani

Social media platforms have become central to global communication, yet they also facilitate the spread of hate speech. For underrepresented dialects like Levantine Arabic, detecting hate speech presents unique cultural, ethical, and…

Computation and Language · Computer Science 2024-12-17 Ahmed Haj Ahmed , Rui-Jie Yew , Xerxes Minocher , Suresh Venkatasubramanian

Entity coreference resolution is an important research problem with many applications, including information extraction and question answering. Coreference resolution for English has been studied extensively. However, there is relatively…

Computation and Language · Computer Science 2023-01-24 Tuan Manh Lai , Heng Ji

Large scale pre-training models have been widely used in named entity recognition (NER) tasks. However, model ensemble through parameter averaging or voting can not give full play to the differentiation advantages of different models,…

Computation and Language · Computer Science 2022-05-31 Changyu Hou , Jun Wang , Yixuan Qiao , Peng Jiang , Peng Gao , Guotong Xie , Qizhi Lin , Xiaopeng Wang , Xiandi Jiang , Benqi Wang , Qifeng Xiao