中文
相关论文

相关论文: Manually Annotated Spelling Error Corpus for Amhar…

200 篇论文

Natural language processing technology has rapidly improved automated grammatical error correction tasks, and the community begins to explore document-level revision as one of the next challenges. To go beyond sentence-level automated…

计算与语言 · 计算机科学 2022-05-24 Masato Mita , Keisuke Sakaguchi , Masato Hagiwara , Tomoya Mizumoto , Jun Suzuki , Kentaro Inui

Wikipedia articles (content pages) are commonly used corpora in Natural Language Processing (NLP) research, especially in low-resource languages other than English. Yet, a few research studies have studied the three Arabic Wikipedia…

计算与语言 · 计算机科学 2024-04-02 Saied Alshahrani , Hesham Haroon , Ali Elfilali , Mariama Njie , Jeanna Matthews

Homophones present a significant challenge to authors in any languages due to their similarities of pronunciations but different meanings and spellings. This issue is particularly pronounced in the Khmer language, rich in homophones due to…

计算与语言 · 计算机科学 2024-11-19 Seanghort Born , Madeth May , Claudine Piau-Toffolon , Sébastien Iksal

Data annotation is an important but time-consuming and costly procedure. To sort a text into two classes, the very first thing we need is a good annotation guideline, establishing what is required to qualify for each class. In the…

计算与语言 · 计算机科学 2018-08-16 Imane Guellil , Ahsan Adeel , Faical Azouaou , Amir Hussain

Yor\`ub\'a is a widely spoken West African language with a writing system rich in tonal and orthographic diacritics. With very few exceptions, diacritics are omitted from electronic texts, due to limited device and application support.…

计算与语言 · 计算机科学 2018-10-31 Iroro Orife

Resources for Grammatical Error Correction (GEC) in non-English languages are scarce, while available spellcheckers in these languages are mostly limited to simple corrections and rules. In this paper we introduce a first GEC corpus for…

计算与语言 · 计算机科学 2026-04-28 Teodor-Mihai Cotet , Stefan Ruseti , Mihai Dascalu

Grammar checking is the task of detection and correction of grammatical errors in the text. English is the dominating language in the field of science and technology. Therefore, the non-native English speakers must be able to use correct…

计算与语言 · 计算机科学 2018-04-03 Madhvi Soni , Jitendra Singh Thakur

This article presents the application of the Universal Named Entity framework to generate automatically annotated corpora. By using a workflow that extracts Wikipedia data and meta-data and DBpedia information, we generated an English…

计算与语言 · 计算机科学 2022-12-15 Diego Alves , Gaurish Thakkar , Marko Tadić

Manually annotated datasets are crucial for training and evaluating Natural Language Processing models. However, recent work has discovered that even widely-used benchmark datasets contain a substantial number of erroneous annotations. This…

计算与语言 · 计算机科学 2023-06-01 Leon Weber , Barbara Plank

In the field of natural language processing, correction of performance assessment for chance agreement plays a crucial role in evaluating the reliability of annotations. However, there is a notable dearth of research focusing on chance…

计算与语言 · 计算机科学 2024-11-18 Diya Li , Carolyn Rosé , Ao Yuan , Chunxiao Zhou

Recent research in natural language processing (NLP) has achieved impressive performance in tasks such as machine translation (MT), news classification, and question-answering in high-resource languages. However, the performance of MT…

计算与语言 · 计算机科学 2024-03-29 Atnafu Lambebo Tonja , Olga Kolesnikova , Alexander Gelbukh , Jugal Kalita

High-quality Machine Translation (MT) evaluation relies heavily on human judgments. Comprehensive error classification methods, such as Multidimensional Quality Metrics (MQM), are expensive as they are time-consuming and can only be done by…

In this work, we introduce the construction of a machine translation (MT) assisted and human-in-the-loop multilingual parallel corpus with annotations of multi-word expressions (MWEs), named AlphaMWE. The MWEs include verbal MWEs (vMWEs)…

计算与语言 · 计算机科学 2025-12-23 Lifeng Han , Najet Hadj Mohamed , Malak Rassem , Gareth Jones , Alan Smeaton , Goran Nenadic

The paper presents methods for evaluating the accuracy of alignments between transcriptions and audio recordings. The methods have been applied to the Spoken British National Corpus, which is an extensive and varied corpus of natural…

声音 · 计算机科学 2011-01-11 Ladan Baghai-Ravary , Sergio Grau , Greg Kochanski

Text Summarization is the task of condensing long text into just a handful of sentences. Many approaches have been proposed for this task, some of the very first were building statistical models (Extractive Methods) capable of selecting…

计算与语言 · 计算机科学 2020-04-02 Amr M. Zaki , Mahmoud I. Khalil , Hazem M. Abbas

In this paper, we present a quantitative evaluation of differences between alternative translations in a large recently released Finnish paraphrase corpus focusing in particular on non-trivial variation in translation. We combine a series…

计算与语言 · 计算机科学 2021-05-07 Li-Hsin Chang , Sampo Pyysalo , Jenna Kanerva , Filip Ginter

Annotated data is an essential ingredient in natural language processing for training and evaluating machine learning models. It is therefore very desirable for the annotations to be of high quality. Recent work, however, has shown that…

计算与语言 · 计算机科学 2022-09-27 Jan-Christoph Klie , Bonnie Webber , Iryna Gurevych

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus, we collected data…

计算与语言 · 计算机科学 2018-08-24 Wajdi Zaghouani , Anis Charfi

We have built SinSpell, a comprehensive spelling checker for the Sinhala language which is spoken by over 16 million people, mainly in Sri Lanka. However, until recently, Sinhala had no spelling checker with acceptable coverage. Sinspell is…

计算与语言 · 计算机科学 2021-07-08 Upuli Liyanapathirana , Kaumini Gunasinghe , Gihan Dias

We introduce a large and diverse Czech corpus annotated for grammatical error correction (GEC) with the aim to contribute to the still scarce data resources in this domain for languages other than English. The Grammar Error Correction…

计算与语言 · 计算机科学 2022-04-22 Jakub Náplava , Milan Straka , Jana Straková , Alexandr Rosen