中文
相关论文

相关论文: The SAMER Arabic Text Simplification Corpus

200 篇论文

Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10000 Arabic tweets and…

计算与语言 · 计算机科学 2025-05-26 Wajdi Zaghouani , Md. Rafiul Biswas

In academia, plagiarism is certainly not an emerging concern, but it became of a greater magnitude with the popularisation of the Internet and the ease of access to a worldwide source of content, rendering human-only intervention…

计算与语言 · 计算机科学 2022-01-11 Mehdi Abdelhamid , Faical Azouaou , Sofiane Batata

Summarizing texts is not a straightforward task. Before even considering text summarization, one should determine what kind of summary is expected. How much should the information be compressed? Is it relevant to reformulate or should the…

计算与语言 · 计算机科学 2020-07-16 Paul Tardy , David Janiszek , Yannick Estève , Vincent Nguyen

Automated Compliance Checking (ACC) systems aim to semantically parse building regulations to a set of rules. However, semantic parsing is known to be hard and requires large amounts of training data. The complexity of creating such…

计算与语言 · 计算机科学 2021-10-05 Ruben Kruiper , Ioannis Konstas , Alasdair Gray , Farhad Sadeghineko , Richard Watson , Bimal Kumar

In this work, we introduce the construction of a machine translation (MT) assisted and human-in-the-loop multilingual parallel corpus with annotations of multi-word expressions (MWEs), named AlphaMWE. The MWEs include verbal MWEs (vMWEs)…

计算与语言 · 计算机科学 2025-12-23 Lifeng Han , Najet Hadj Mohamed , Malak Rassem , Gareth Jones , Alan Smeaton , Goran Nenadic

The increasing volume of textual data poses challenges in reading and comprehending large documents, particularly for scholars who need to extract useful information from research articles. Automatic text summarization has emerged as a…

计算与语言 · 计算机科学 2025-03-14 Samira Zangooei , Amirhossein Darmani , Hossein Farahmand Nezhad , Laya Mahmoudi

Materials science literature contains millions of materials synthesis procedures described in unstructured natural language text. Large-scale analysis of these synthesis procedures would facilitate deeper scientific understanding of…

The complexities of Arabic language in morphology, orthography and dialects makes sentiment analysis for Arabic more challenging. Also, text feature extraction from short messages like tweets, in order to gauge the sentiment, makes this…

计算与语言 · 计算机科学 2018-10-17 Abdulaziz M. Alayba , Vasile Palade , Matthew England , Rahat Iqbal

Calligraphy is an essential part of the Arabic heritage and culture. It has been used in the past for the decoration of houses and mosques. Usually, such calligraphy is designed manually by experts with aesthetic insights. In the past few…

计算与语言 · 计算机科学 2021-06-28 Zaid Alyafeai , Maged S. Al-shaibani , Mustafa Ghaleb , Yousif Ahmed Al-Wajih

The ambition of a character recognition system is to transform a text document typed on paper into a digital format that can be manipulated by word processor software Unlike other languages, Arabic has unique features, while other language…

计算与语言 · 计算机科学 2010-06-15 A. A Zaidan , B. B Zaidan , Hamid. A. Jalab , Hamdan. O. Alanazi , Rami Alnaqeib

At present, Text-to-speech (TTS) systems that are trained with high-quality transcribed speech data using end-to-end neural models can generate speech that is intelligible, natural, and closely resembles human speech. These models are…

计算与语言 · 计算机科学 2023-03-02 Ajinkya Kulkarni , Atharva Kulkarni , Sara Abedalmonem Mohammad Shatnawi , Hanan Aldarmaki

In this paper, we address the problems of Arabic Text Classification and stemming using Transducers and Rational Kernels. We introduce a new stemming technique based on the use of Arabic patterns (Pattern Based Stemmer). Patterns are…

计算与语言 · 计算机科学 2015-02-27 Attia Nehar , Djelloul Ziadi , Hadda Cherroun

Predicting which words are considered hard to understand for a given target population is a vital step in many NLP applications such as text simplification. This task is commonly referred to as Complex Word Identification (CWI). With a few…

计算与语言 · 计算机科学 2020-06-12 Matthew Shardlow , Michael Cooper , Marcos Zampieri

Over the past years, interest in discourse analysis and discourse parsing has steadily grown, and many discourse-annotated corpora and, as a result, discourse parsers have been built. In this paper, we present a discourse-annotated corpus…

计算与语言 · 计算机科学 2021-06-29 Sara Shahmohammadi , Hadi Veisi , Ali Darzi

We present a freely available, genre-balanced English web corpus totaling 4M tokens and featuring a large number of high-quality automatic annotation layers, including dependency trees, non-named entity annotations, coreference resolution,…

计算与语言 · 计算机科学 2020-06-19 Luke Gessler , Siyao Peng , Yang Liu , Yilun Zhu , Shabnam Behzad , Amir Zeldes

This paper presents the Arabic Women and Society Corpus, a ten year collection of 252,487 public Arabic Facebook posts related to women's empowerment and social wellbeing. The corpus was collected from 51,660 pages across 77 countries…

计算与语言 · 计算机科学 2026-05-22 Wajdi Zaghouani , Mabrouka Bessghaier , MD. Rafiul Biswas , Shimaa Amer Ibrahim

This paper focuses on detecting propagandistic spans and persuasion techniques in Arabic text from tweets and news paragraphs. Each entry in the dataset contains a text sample and corresponding labels that indicate the start and end…

计算与语言 · 计算机科学 2024-08-09 Md Rafiul Biswas , Zubair Shah , Wajdi Zaghouani

In this paper, we introduce SaudiBERT, a monodialect Arabic language model pretrained exclusively on Saudi dialectal text. To demonstrate the model's effectiveness, we compared SaudiBERT with six different multidialect Arabic language…

计算与语言 · 计算机科学 2024-05-13 Faisal Qarah

As large language models (LLMs) grow and develop, so do their data demands. This is especially true for multilingual LLMs, where the scarcity of high-quality and readily available data online has led to a multitude of synthetic dataset…

计算与语言 · 计算机科学 2024-11-12 Sultan Alrashed , Dmitrii Khizbullin , David R. Pugh

This paper describes our submission to SemEval-2022 Task 6 on sarcasm detection and its five subtasks for English and Arabic. Sarcasm conveys a meaning which contradicts the literal meaning, and it is mainly found on social networks. It has…

计算与语言 · 计算机科学 2022-03-09 Shubham Kumar Nigam , Mosab Shaheen