English
Related papers

Related papers: The Grammar and Syntax Based Corpus Analysis Tool …

200 papers

This work aims to provide an overview on the open-source multilanguage tool called StyloMetrix. It offers stylometric text representations that cover various aspects of grammar, syntax and lexicon. StyloMetrix covers four languages: Polish…

Computation and Language · Computer Science 2023-09-25 Inez Okulska , Daria Stetsenko , Anna Kołos , Agnieszka Karlińska , Kinga Głąbińska , Adam Nowakowski

We introduce Spivavtor, a dataset, and instruction-tuned models for text editing focused on the Ukrainian language. Spivavtor is the Ukrainian-focused adaptation of the English-only CoEdIT model. Similar to CoEdIT, Spivavtor performs text…

Computation and Language · Computer Science 2024-04-30 Aman Saini , Artem Chernodub , Vipul Raheja , Vivek Kulkarni

Due to the growing role of the SEO technologies, it is necessary to perform an automated analysis of the article's quality. Such approach helps both to return the most intelligible pages for the user's query and to raise the web sites…

Computation and Language · Computer Science 2020-11-03 S. D. Pogorilyy , A. A. Kramov

The paper explores stylometry as a method to distinguish between texts created by Large Language Models (LLMs) and humans, addressing issues of model attribution, intellectual property, and ethical AI use. Stylometry has been used…

Computation and Language · Computer Science 2025-07-25 Karol Przystalski , Jan K. Argasiński , Iwona Grabska-Gradzińska , Jeremi K. Ochab

As the usage of large language models for problems outside of simple text understanding or generation increases, assessing their abilities and limitations becomes crucial. While significant progress has been made in this area over the last…

Computation and Language · Computer Science 2025-01-14 Mykyta Syromiatnikov , Victoria Ruvinskaya , Anastasiya Troynina

When solving tasks in the field of natural language processing, we sometimes need dictionary tools, such as lexicons, word form dictionaries or knowledge bases. However, the availability of dictionary data is insufficient in many languages,…

Computation and Language · Computer Science 2025-12-02 Miroslav Blšták

While the evaluation of multimodal English-centric models is an active area of research with numerous benchmarks, there is a profound lack of benchmarks or evaluation suites for low- and mid-resource languages. We introduce ZNO-Vision, a…

Computation and Language · Computer Science 2024-11-25 Yurii Paniv , Artur Kiulian , Dmytro Chaplynskyi , Mykola Khandoga , Anton Polishko , Tetiana Bas , Guillermo Gabrielli

pymorphy2 is a morphological analyzer and generator for Russian and Ukrainian languages. It uses large efficiently encoded lexi- cons built from OpenCorpora and LanguageTool data. A set of linguistically motivated rules is developed to…

Computation and Language · Computer Science 2015-03-26 Mikhail Korobov

The task of toxicity detection is still a relevant task, especially in the context of safe and fair LMs development. Nevertheless, labeled binary toxicity classification corpora are not available for all languages, which is understandable…

Computation and Language · Computer Science 2024-04-30 Daryna Dementieva , Valeriia Khylenko , Nikolay Babakov , Georg Groh

The ability to compare the semantic similarity between text corpora is important in a variety of natural language processing applications. However, standard methods for evaluating these metrics have yet to be established. We propose a set…

Computation and Language · Computer Science 2022-11-30 George Kour , Samuel Ackerman , Orna Raz , Eitan Farchi , Boaz Carmeli , Ateret Anaby-Tavor

Metamodel-based DSL development in language workbenches like Xtext allows language engineers to focus more on metamodels and domain concepts rather than grammar details. However, the grammar generated from metamodels often requires manual…

Software Engineering · Computer Science 2023-09-11 Weixing Zhang , Jan-Philipp Steghöfer , Regina Hebig , Daniel Strüber

We present a corpus professionally annotated for grammatical error correction (GEC) and fluency edits in the Ukrainian language. To the best of our knowledge, this is the first GEC corpus for the Ukrainian language. We collected texts with…

Computation and Language · Computer Science 2022-11-09 Oleksiy Syvokon , Olena Nahorna

Supervised text models are a valuable tool for political scientists but present several obstacles to their use, including the expense of hand-labeling documents, the difficulty of retrieving rare relevant documents for annotation, and…

Computation and Language · Computer Science 2025-06-18 Andrew Halterman

To build large language models for Ukrainian we need to expand our corpora with large amounts of new algorithmic tasks expressed in natural language. Examples of task performance expressed in English are abundant, so with a high-quality…

Computation and Language · Computer Science 2024-07-15 Yurii Paniv , Dmytro Chaplynskyi , Nikita Trynus , Volodymyr Kyrylov

This submission to the binary AI detection task is based on a modular stylometric pipeline, where: public spaCy models are used for text preprocessing (including tokenisation, named entity recognition, dependency parsing, part-of-speech…

Computation and Language · Computer Science 2025-07-17 Jeremi K. Ochab , Mateusz Matias , Tymoteusz Boba , Tomasz Walkowiak

The increasing sophistication of AI-generated texts highlights the urgent need for accurate and transparent detection tools, especially in educational settings, where verifying authorship is essential. Existing literature has demonstrated…

Computation and Language · Computer Science 2025-05-06 Chidimma Opara

In this work, we introduce skLEP, the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. We have compiled skLEP to encompass nine diverse tasks that span token-level,…

Stylometry, the science of inferring characteristics of the author from the characteristics of documents written by that author, is a problem with a long history and belongs to the core task of Text categorization that involves authorship…

Computation and Language · Computer Science 2012-10-16 Tanmoy Chakraborty , Sivaji Bandyopadhyay

The syntactic behaviour of texts can highly vary depending on their contexts (e.g. author, genre, etc.). From the standpoint of stylometry, it can be helpful to objectively measure this behaviour. In this paper, we discuss how coalgebras…

Computation and Language · Computer Science 2021-08-10 Joël A. Doat

Folktales are linguistically very rich and culturally significant in understanding the source language. Historically, only human translation has been used for translating folklore. Therefore, the number of translated texts is very sparse,…

Computation and Language · Computer Science 2024-10-15 Olena Burda-Lassen
‹ Prev 1 2 3 10 Next ›