中文
相关论文

相关论文: The Grammar and Syntax Based Corpus Analysis Tool …

200 篇论文

This work aims to provide an overview on the open-source multilanguage tool called StyloMetrix. It offers stylometric text representations that cover various aspects of grammar, syntax and lexicon. StyloMetrix covers four languages: Polish…

计算与语言 · 计算机科学 2023-09-25 Inez Okulska , Daria Stetsenko , Anna Kołos , Agnieszka Karlińska , Kinga Głąbińska , Adam Nowakowski

We introduce Spivavtor, a dataset, and instruction-tuned models for text editing focused on the Ukrainian language. Spivavtor is the Ukrainian-focused adaptation of the English-only CoEdIT model. Similar to CoEdIT, Spivavtor performs text…

计算与语言 · 计算机科学 2024-04-30 Aman Saini , Artem Chernodub , Vipul Raheja , Vivek Kulkarni

Due to the growing role of the SEO technologies, it is necessary to perform an automated analysis of the article's quality. Such approach helps both to return the most intelligible pages for the user's query and to raise the web sites…

计算与语言 · 计算机科学 2020-11-03 S. D. Pogorilyy , A. A. Kramov

The paper explores stylometry as a method to distinguish between texts created by Large Language Models (LLMs) and humans, addressing issues of model attribution, intellectual property, and ethical AI use. Stylometry has been used…

计算与语言 · 计算机科学 2025-07-25 Karol Przystalski , Jan K. Argasiński , Iwona Grabska-Gradzińska , Jeremi K. Ochab

As the usage of large language models for problems outside of simple text understanding or generation increases, assessing their abilities and limitations becomes crucial. While significant progress has been made in this area over the last…

计算与语言 · 计算机科学 2025-01-14 Mykyta Syromiatnikov , Victoria Ruvinskaya , Anastasiya Troynina

When solving tasks in the field of natural language processing, we sometimes need dictionary tools, such as lexicons, word form dictionaries or knowledge bases. However, the availability of dictionary data is insufficient in many languages,…

计算与语言 · 计算机科学 2025-12-02 Miroslav Blšták

While the evaluation of multimodal English-centric models is an active area of research with numerous benchmarks, there is a profound lack of benchmarks or evaluation suites for low- and mid-resource languages. We introduce ZNO-Vision, a…

pymorphy2 is a morphological analyzer and generator for Russian and Ukrainian languages. It uses large efficiently encoded lexi- cons built from OpenCorpora and LanguageTool data. A set of linguistically motivated rules is developed to…

计算与语言 · 计算机科学 2015-03-26 Mikhail Korobov

The task of toxicity detection is still a relevant task, especially in the context of safe and fair LMs development. Nevertheless, labeled binary toxicity classification corpora are not available for all languages, which is understandable…

计算与语言 · 计算机科学 2024-04-30 Daryna Dementieva , Valeriia Khylenko , Nikolay Babakov , Georg Groh

The ability to compare the semantic similarity between text corpora is important in a variety of natural language processing applications. However, standard methods for evaluating these metrics have yet to be established. We propose a set…

计算与语言 · 计算机科学 2022-11-30 George Kour , Samuel Ackerman , Orna Raz , Eitan Farchi , Boaz Carmeli , Ateret Anaby-Tavor

Metamodel-based DSL development in language workbenches like Xtext allows language engineers to focus more on metamodels and domain concepts rather than grammar details. However, the grammar generated from metamodels often requires manual…

软件工程 · 计算机科学 2023-09-11 Weixing Zhang , Jan-Philipp Steghöfer , Regina Hebig , Daniel Strüber

We present a corpus professionally annotated for grammatical error correction (GEC) and fluency edits in the Ukrainian language. To the best of our knowledge, this is the first GEC corpus for the Ukrainian language. We collected texts with…

计算与语言 · 计算机科学 2022-11-09 Oleksiy Syvokon , Olena Nahorna

Supervised text models are a valuable tool for political scientists but present several obstacles to their use, including the expense of hand-labeling documents, the difficulty of retrieving rare relevant documents for annotation, and…

计算与语言 · 计算机科学 2025-06-18 Andrew Halterman

To build large language models for Ukrainian we need to expand our corpora with large amounts of new algorithmic tasks expressed in natural language. Examples of task performance expressed in English are abundant, so with a high-quality…

计算与语言 · 计算机科学 2024-07-15 Yurii Paniv , Dmytro Chaplynskyi , Nikita Trynus , Volodymyr Kyrylov

This submission to the binary AI detection task is based on a modular stylometric pipeline, where: public spaCy models are used for text preprocessing (including tokenisation, named entity recognition, dependency parsing, part-of-speech…

计算与语言 · 计算机科学 2025-07-17 Jeremi K. Ochab , Mateusz Matias , Tymoteusz Boba , Tomasz Walkowiak

The increasing sophistication of AI-generated texts highlights the urgent need for accurate and transparent detection tools, especially in educational settings, where verifying authorship is essential. Existing literature has demonstrated…

计算与语言 · 计算机科学 2025-05-06 Chidimma Opara

In this work, we introduce skLEP, the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. We have compiled skLEP to encompass nine diverse tasks that span token-level,…

Stylometry, the science of inferring characteristics of the author from the characteristics of documents written by that author, is a problem with a long history and belongs to the core task of Text categorization that involves authorship…

计算与语言 · 计算机科学 2012-10-16 Tanmoy Chakraborty , Sivaji Bandyopadhyay

The syntactic behaviour of texts can highly vary depending on their contexts (e.g. author, genre, etc.). From the standpoint of stylometry, it can be helpful to objectively measure this behaviour. In this paper, we discuss how coalgebras…

计算与语言 · 计算机科学 2021-08-10 Joël A. Doat

Folktales are linguistically very rich and culturally significant in understanding the source language. Historically, only human translation has been used for translating folklore. Therefore, the number of translated texts is very sparse,…

计算与语言 · 计算机科学 2024-10-15 Olena Burda-Lassen
‹ 上一页 1 2 3 10 下一页 ›