English
Related papers

Related papers: Manually Annotated Spelling Error Corpus for Amhar…

200 papers

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowd-sourcing initiative, encompassing a diverse range…

Computation and Language · Computer Science 2025-02-04 Turi Abu , Ying Shi , Thomas Fang Zheng , Dong Wang

Most languages, especially in Africa, have fewer or no established part-of-speech (POS) tagged corpus. However, POS tagged corpus is essential for natural language processing (NLP) to support advanced researches such as machine translation,…

Computation and Language · Computer Science 2019-03-14 Onyenwe Ikechukwu E , Onyedinma Ebele G , Aniegwu Godwin E , Ezeani Ignatius M

This article investigates the use of Transformation-Based Error-Driven learning for resolving part-of-speech ambiguity in the Greek language. The aim is not only to study the performance, but also to examine its dependence on different…

Computation and Language · Computer Science 2007-05-23 G. Petasis , G. Paliouras , V. Karkaletsis , C. D. Spyropoulos , I. Androutsopoulos

Automatic spelling and grammatical correction systems are one of the most widely used tools within natural language applications. In this thesis, we assume the task of error correction as a type of monolingual machine translation where the…

Computation and Language · Computer Science 2018-10-02 Sina Ahmadi

This paper introduces a spelling correction system which integrates seamlessly with morphological analysis using a multi-tape formalism. Handling of various Semitic error problems is illustrated, with reference to Arabic and Syriac…

cmp-lg · Computer Science 2008-02-03 Tanya Bowden , George Anton Kiraz

Annotating a multilingual code-switched corpus is a painstaking process requiring specialist linguistic expertise. This is partly due to the large number of language combinations that may appear within and across utterances, which might…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-18 Geoffrey Frost , Emily Morris , Joshua Jansen van Vüren , Thomas Niesler

Emotion is a crucial phenomenon in the functioning of human beings in society. However, it remains a widely open subject, particularly in its textual manifestations. This paper examines an industrial corpus manually annotated following an…

Computation and Language · Computer Science 2025-09-03 Jonas Noblet

This paper reports on the preliminary phase of our ongoing research towards developing an intelligent tutoring environment for Turkish grammar. One of the components of this environment is a corpus search tool which, among other aspects of…

cmp-lg · Computer Science 2016-08-31 H. Altay Guvenir , Kemal Oflazer

The great amount of information that can be stored in electronic media is growing up daily. Many of them is got mainly by typing, such as the huge of information obtained from web 2.0 sites; or scaned and processing by an Optical Character…

Computation and Language · Computer Science 2021-12-06 Wulfrano A. Luna-Ramírez , Carlos R. Jaimez-González

While there exist strong benchmark datasets for grammatical error correction (GEC), high-quality annotated spoken datasets for Spoken GEC (SGEC) are still under-resourced. In this paper, we propose a fully automated method to generate…

Computation and Language · Computer Science 2025-07-28 Penny Karanasou , Mengjie Qian , Stefano Bannò , Mark J. F. Gales , Kate M. Knill

We present a new parallel corpus, JHU FLuency-Extended GUG corpus (JFLEG) for developing and evaluating grammatical error correction (GEC). Unlike other corpora, it represents a broad range of language proficiency levels and uses holistic…

Computation and Language · Computer Science 2017-02-15 Courtney Napoles , Keisuke Sakaguchi , Joel Tetreault

In a conventional Speech emotion recognition (SER) task, a classifier for a given language is trained on a pre-existing dataset for that same language. However, where training data for a language does not exist, data from other languages…

In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently a lack of available…

Computation and Language · Computer Science 2019-12-03 Nelda Kote , Marenglen Biba , Jenna Kanerva , Samuel Rönnqvist , Filip Ginter

SALMA, the first Arabic sense-annotated corpus, consists of ~34K tokens, which are all sense-annotated. The corpus is annotated using two different sense inventories simultaneously (Modern and Ghani). SALMA novelty lies in how tokens and…

Computation and Language · Computer Science 2023-10-31 Mustafa Jarrar , Sanad Malaysha , Tymaa Hammouda , Mohammed Khalilia

Spelling correction is the task of identifying spelling mistakes, typos, and grammatical mistakes in a given text and correcting them according to their context and grammatical structure. This work introduces "AraSpell," a framework for…

Computation and Language · Computer Science 2024-05-14 Mahmoud Salhab , Faisal Abu-Khzam

We present a novel method of performing spelling correction on short input strings, such as search queries or individual words. At its core lies a procedure for generating artificial typos which closely follow the error patterns manifested…

Computation and Language · Computer Science 2021-05-14 Alex Kuznetsov , Hector Urdiales

In this paper, we present the annotation pipeline and the guidelines we wrote as part of an effort to create a large manually annotated Arabic author profiling dataset from various social media sources covering 16 Arabic countries and 11…

Computation and Language · Computer Science 2018-08-24 Wajdi Zaghouani , Anis Charfi

When learning grammar of the new language, a teacher should routinely check student's exercises for grammatical correctness. The paper describes a method of automatically detecting and reporting grammar mistakes, regarding an order of…

Computation and Language · Computer Science 2013-01-14 Oleg Sychev , Dmitry Mamontov

Automated audio captioning (AAC) is an important cross-modality translation task, aiming at generating descriptions for audio clips. However, captions generated by previous AAC models have faced ``false-repetition'' errors due to the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-21 Hanxue Zhang , Zeyu Xie , Xuenan Xu , Mengyue Wu , Kai Yu

This study illustrates how incorporating feedback-oriented annotations into the scoring pipeline can enhance the accuracy of automated essay scoring (AES). This approach is demonstrated with the Persuasive Essays for Rating, Selecting, and…

Computation and Language · Computer Science 2025-09-03 Christopher Ormerod