中文
相关论文

相关论文: The SAMER Arabic Text Simplification Corpus

200 篇论文

Text simplification is crucial for improving accessibility and comprehension for English as a Second Language (ESL) learners. This study goes a step further and aims to facilitate ESL learners' language acquisition by simplification.…

计算与语言 · 计算机科学 2025-02-18 Guanlin Li , Yuki Arase , Noel Crespi

In this paper, we introduce MADARi, a joint morphological annotation and spelling correction system for texts in Standard and Dialectal Arabic. The MADARi framework provides intuitive interfaces for annotating text and managing the…

计算与语言 · 计算机科学 2018-08-28 Ossama Obeid , Salam Khalifa , Nizar Habash , Houda Bouamor , Wajdi Zaghouani , Kemal Oflazer

Gender bias in natural language processing (NLP) applications, particularly machine translation, has been receiving increasing attention. Much of the research on this issue has focused on mitigating gender bias in English NLP models and…

计算与语言 · 计算机科学 2021-10-19 Bashar Alhafni , Nizar Habash , Houda Bouamor

We present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This…

计算与语言 · 计算机科学 2024-10-17 Tarek Naous , Michael J. Ryan , Anton Lavrouk , Mohit Chandra , Wei Xu

We introduce S2ORC, a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for…

计算与语言 · 计算机科学 2020-07-08 Kyle Lo , Lucy Lu Wang , Mark Neumann , Rodney Kinney , Dan S. Weld

Arabic is a widely-spoken language with a long and rich history, but existing corpora and language technology focus mostly on modern Arabic and its varieties. Therefore, studying the history of the language has so far been mostly limited to…

计算与语言 · 计算机科学 2018-09-12 Yonatan Belinkov , Alexander Magidow , Alberto Barrón-Cedeño , Avi Shmidman , Maxim Romanov

This paper describes the acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. It will be helpful for machine translation of a low-resource language, Amharic. We freely released the corpus for…

计算与语言 · 计算机科学 2022-05-04 Andargachew Mekonnen Gezmu , Andreas Nürnberger , Tesfaye Bayu Bati

Matching texts in highly inflected languages such as Arabic by simple stemming strategy is unlikely to perform well. In this paper, we present a strategy for automatic text matching technique for for inflectional languages, using Arabic as…

计算与语言 · 计算机科学 2014-03-25 Tarek El-Shishtawy , Fatma El-Ghannam

We present hinglishNorm -- a human annotated corpus of Hindi-English code-mixed sentences for text normalization task. Each sentence in the corpus is aligned to its corresponding human annotated normalized form. To the best of our…

计算与语言 · 计算机科学 2020-10-20 Piyush Makhija , Ankit Kumar , Anuj Gupta

Training learnable metrics using modern language models has recently emerged as a promising method for the automatic evaluation of machine translation. However, existing human evaluation datasets for text simplification have limited…

计算与语言 · 计算机科学 2023-07-11 Mounica Maddela , Yao Dou , David Heineman , Wei Xu

This paper proposes a methodology to prepare corpora in Arabic language from online social network (OSN) and review site for Sentiment Analysis (SA) task. The paper also proposes a methodology for generating a stopword list from the…

计算与语言 · 计算机科学 2014-10-07 Walaa Medhat , Ahmed H. Yousef , Hoda Korashy

This paper presents a manually annotated spelling error corpus for Amharic, lingua franca in Ethiopia. The corpus is designed to be used for the evaluation of spelling error detection and correction. The misspellings are tagged as non-word…

Determining the readability of a text is the first step to its simplification. In this paper, we present a readability analysis tool capable of analyzing text written in the Bengali language to provide in-depth information on its…

计算与语言 · 计算机科学 2020-12-15 Susmoy Chakraborty , Mir Tafseer Nayeem , Wasi Uddin Ahmad

The rise of large language models (LLMs) has transformed numerous natural language processing (NLP) tasks, yet their performance in low and mid-resource languages, such as Farsi, still lags behind resource-rich languages like English. To…

计算与语言 · 计算机科学 2024-12-24 Sadra Sabouri , Elnaz Rahmati , Soroush Gooran , Hossein Sameti

The goal of this paper is to present a model of children's semantic memory, which is based on a corpus reproducing the kinds of texts children are exposed to. After presenting the literature in the development of the semantic memory, a…

计算与语言 · 计算机科学 2008-12-18 Guy Denhière , Benoît Lemaire , Cédrick Bellissens , Sandra Jhean

We present a unified benchmark for mispronunciation detection in Modern Standard Arabic (MSA) using Qur'anic recitation as a case study. Our approach lays the groundwork for advancing Arabic pronunciation assessment by providing a…

This paper presents Nabra, a corpora of Syrian Arabic dialects with morphological annotations. A team of Syrian natives collected more than 6K sentences containing about 60K words from several sources including social media posts, scripts…

计算与语言 · 计算机科学 2023-10-27 Amal Nayouf , Tymaa Hammouda , Mustafa Jarrar , Fadi Zaraket , Mohamad-Bassam Kurdy

A prototype system for the transliteration of diacritics-less Arabic manuscripts at the sub-word or part of Arabic word (PAW) level is developed. The system is able to read sub-words of the input manuscript using a set of skeleton-based…

计算机视觉与模式识别 · 计算机科学 2013-06-27 Reza Farrahi Moghaddam , Mohamed Cheriet , Thomas Milo , Robert Wisnovsky

Text categorization is the process of grouping documents into categories based on their contents. This process is important to make information retrieval easier, and it became more important due to the huge textual information available…

信息检索 · 计算机科学 2015-01-08 Ashraf Odeh , Aymen Abu-Errub , Qusai Shambour , Nidal Turab

We introduced the contemporary Amharic corpus, which is automatically tagged for morpho-syntactic information. Texts are collected from 25,199 documents from different domains and about 24 million orthographic words are tokenized. Since it…

计算与语言 · 计算机科学 2021-06-15 Andargachew Mekonnen Gezmu , Binyam Ephrem Seyoum , Michael Gasser , Andreas Nürnberger