中文
相关论文

相关论文: Automatic Romanization of Arabic Bibliographic Rec…

200 篇论文

We observe a recent behaviour on social media, in which users intentionally remove consonantal dots from Arabic letters, in order to bypass content-classification algorithms. Content classification is typically done by fine-tuning…

计算与语言 · 计算机科学 2021-11-19 Aviad Rom , Kfir Bar

This article describes the constitution process of the first morpho-syntactically annotated Tunisian Arabish Corpus (TArC). Arabish, also known as Arabizi, is a spontaneous coding of Arabic dialects in Latin characters and arithmographs…

计算与语言 · 计算机科学 2020-11-11 Elisa Gugliotta , Marco Dinarelli

The process of language learning involves the mastery of countless tasks: making the constituent sounds of the language being learned, learning the grammatical patterns, and acquiring the requisite vocabulary for reception and production.…

人工智能 · 计算机科学 2015-04-09 Hayat Alrefaie , Allan Ramsay

In this paper, we address the problems of Arabic Text Classification and stemming using Transducers and Rational Kernels. We introduce a new stemming technique based on the use of Arabic patterns (Pattern Based Stemmer). Patterns are…

计算与语言 · 计算机科学 2015-02-27 Attia Nehar , Djelloul Ziadi , Hadda Cherroun

Arabic handwriting is a consonantal and cursive writing. The analysis of Arabic script is further complicated due to obligatory dots/strokes that are placed above or below most letters and usually written delayed in order. Due to…

计算机视觉与模式识别 · 计算机科学 2015-10-20 Ibrahim Abdelaziz , Sherif Abdou , Hassanin Al-Barhamtoshy

This paper describes the acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. It will be helpful for machine translation of a low-resource language, Amharic. We freely released the corpus for…

计算与语言 · 计算机科学 2022-05-04 Andargachew Mekonnen Gezmu , Andreas Nürnberger , Tesfaye Bayu Bati

In this paper, we address the significant gap in Arabic natural language processing (NLP) resources by introducing ArabicaQA, the first large-scale dataset for machine reading comprehension and open-domain question answering in Arabic. This…

Code-switching poses a number of challenges and opportunities for multilingual automatic speech recognition. In this paper, we focus on the question of robust and fair evaluation metrics. To that end, we develop a reference benchmark data…

音频与语音处理 · 电气工程与系统科学 2022-11-30 Injy Hamed , Amir Hussein , Oumnia Chellah , Shammur Chowdhury , Hamdy Mubarak , Sunayana Sitaram , Nizar Habash , Ahmed Ali

Question semantic similarity is a challenging and active research problem that is very useful in many NLP applications, such as detecting duplicate questions in community question answering platforms such as Quora. Arabic is considered to…

计算与语言 · 计算机科学 2019-09-23 Hesham Al-Bataineh , Wael Farhan , Ahmad Mustafa , Haitham Seelawi , Hussein T. Al-Natsheh

The detection of toxic language in the Arabic language has emerged as an active area of research in recent years, and reviewing the existing datasets employed for training the developed solutions has become a pressing need. This paper…

计算与语言 · 计算机科学 2024-01-31 Imene Bensalem , Paolo Rosso , Hanane Zitouni

Labelling of user's utterances to understanding his attends which called Dialogue Act (DA) classification, it is considered the key player for dialogue language understanding layer in automatic dialogue systems. In this paper, we proposed a…

计算与语言 · 计算机科学 2015-09-11 Abdelrahim A Elmadany , Sherif M Abdou , Mervat Gheith

Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces the challenge of being under-resourced. The scarcity of large annotated datasets hampers the development of accurate tools for subjectivity…

计算与语言 · 计算机科学 2026-03-02 Slimane Bellaouar , Attia Nehar , Soumia Souffi , Mounia Bouameur

Despite the growing importance of Arabic as a global language, there is a notable lack of language models pre-trained exclusively on Arabic data. This shortage has led to limited benchmarks available for assessing language model performance…

计算与语言 · 计算机科学 2024-07-02 Shahad Al-Khalifa , Hend Al-Khalifa

As a result of ICT development and the increasingly growing use of the Internet in particular, practices of language teaching and learning are about to evolve significantly. Our study focuses on the Arabic language, and aims to explore and…

其他计算机科学 · 计算机科学 2012-11-06 Hadhemi Achour , Wahiba Ben Abdessalem

The debut of chatGPT and BARD has popularized instruction following text generation using LLMs, where a user can interrogate an LLM using natural language requests and obtain natural language answers that matches their requests. Training…

计算与语言 · 计算机科学 2024-08-13 Abdelrahman El-Sheikh , Ahmed Elmogtaba , Kareem Darwish , Muhammad Elmallah , Ashraf Elneima , Hassan Sawaf

The acceleration in telecommunication needs leads to many groups of research, especially in communication facilitating and Machine Translation fields. While people contact with others having different languages and cultures, they need to…

计算与语言 · 计算机科学 2019-02-07 Neama Abdulaziz Dahan , Fadl Mutaher Ba-Alwi

Modern Standard Arabic (MSA) nominals present many morphological and lexical modeling challenges that have not been consistently addressed previously. This paper attempts to define the space of such challenges, and leverage a recently…

计算与语言 · 计算机科学 2024-02-02 Christian Khairallah , Reham Marzouk , Salam Khalifa , Mayar Nassar , Nizar Habash

The commonsense understanding and validation remains a challenging task in the field of natural language understanding. Therefore, several research papers have been published that studied the capability of proposed systems to evaluate the…

计算与语言 · 计算机科学 2020-08-26 Saja Tawalbeh , Mohammad AL-Smadi

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects alongside Modern…

A common bottleneck for developing machine translation (MT) systems for some language pairs is the lack of direct parallel translation data sets, in general and in certain domains. Alternative solutions such as zero-shot models or pivoting…

计算与语言 · 计算机科学 2019-07-16 Ella Noll , Mai Oudah , Nizar Habash