English
Related papers

Related papers: The Grammar and Syntax Based Corpus Analysis Tool …

200 papers

Large language models (LLMs) are known to exhibit biases in downstream tasks, especially when dealing with sensitive topics such as political discourse, gender identity, ethnic relations, or national stereotypes. Although significant…

Computation and Language · Computer Science 2025-08-18 Martin Pavlíček , Tomáš Filip , Petr Sosík

Language students are most engaged while reading texts at an appropriate difficulty level. However, existing methods of evaluating text difficulty focus mainly on vocabulary and do not prioritize grammatical features, hence they do not work…

Computation and Language · Computer Science 2017-02-17 Shuhan Wang , Erik Andersen

In this article, we describe some discursive segmentation methods as well as a preliminary evaluation of the segmentation quality. Although our experiment were carried for documents in French, we have developed three discursive segmentation…

Computation and Language · Computer Science 2020-06-15 Rémy Saksik , Alejandro Molina-Villegas , Andréa Carneiro Linhares , Juan-Manuel Torres-Moreno

Psychometric measures of ability, attitudes, perceptions, and beliefs are crucial for understanding user behaviors in various contexts including health, security, e-commerce, and finance. Traditionally, psychometric dimensions have been…

Computation and Language · Computer Science 2020-07-28 Ahmed Abbasi , David G. Dobolyi , Richard G. Netemeyer

This paper presents the first Swedish evaluation benchmark for textual semantic similarity. The benchmark is compiled by simply running the English STS-B dataset through the Google machine translation API. This paper discusses potential…

Computation and Language · Computer Science 2020-12-01 Tim Isbister , Magnus Sahlgren

Keyphrase extraction for morphologically rich, low-resource languages remains understudied, largely due to the scarcity of suitable evaluation datasets. We address this gap for Slovak by constructing a dataset of 227,432 scientific…

Computation and Language · Computer Science 2026-03-17 David Števaňák , Marek Šuppa

Methods for learning sentence representations have been actively developed in recent years. However, the lack of pre-trained models and datasets annotated at the sentence level has been a problem for low-resource languages such as Polish…

Computation and Language · Computer Science 2020-01-24 Sławomir Dadas , Michał Perełkiewicz , Rafał Poświata

Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based on Czech historical documents, containing human-defined…

Computation and Language · Computer Science 2026-03-05 Martin Kostelník , Michal Hradiš , Martin Dočekal

This paper describes the speech processing activities conducted at the Polish consortium of the CLARIN project. The purpose of this segment of the project was to develop specific tools that would allow for automatic and semi-automatic…

Computation and Language · Computer Science 2017-06-02 Danijel Koržinek , Krzysztof Marasek , Łukasz Brocki , Krzysztof Wołk

This paper presents Summary Workbench, a new tool for developing and evaluating text summarization models. New models and evaluation measures can be easily integrated as Docker-based plugins, allowing to examine the quality of their…

Computation and Language · Computer Science 2022-10-19 Shahbaz Syed , Dominik Schwabe , Martin Potthast

Despite the extensive amount of labeled datasets in the NLP text classification field, the persistent imbalance in data availability across various languages remains evident. To support further fair development of NLP models, exploring the…

Computation and Language · Computer Science 2025-02-06 Daryna Dementieva , Valeriia Khylenko , Georg Groh

Over the recent years, there has been a growing interest in developing new research evaluation methods that could go beyond the traditional citation-based metrics. This interest is motivated on one side by the wider availability or even…

Digital Libraries · Computer Science 2016-11-17 Drahomira Herrmannova , Petr Knoth

This new research explores the effects of various training methods on a Polish to English Statistical Machine Translation system for medical texts. Various elements of the EMEA parallel text corpora from the OPUS project were used as the…

Computation and Language · Computer Science 2015-09-30 Krzysztof Wołk , Krzysztof Marasek

The lack of automatic evaluation metrics tailored for SignWriting presents a significant obstacle in developing effective transcription and translation models for signed languages. This paper introduces a comprehensive suite of evaluation…

Computation and Language · Computer Science 2024-10-18 Amit Moryossef , Rotem Zilberman , Ohad Langer

Text mining analysis of tweets gathered during Polish presidential election on May 10th, 2015. The project included implementation of engine to retrieve information from Twitter, building document corpora, corpora cleaning, and creating…

Computation and Language · Computer Science 2019-11-05 Karol Chlasta

Generative poetry systems require effective tools for data engineering and automatic evaluation, particularly to assess how well a poem adheres to versification rules, such as the correct alternation of stressed and unstressed syllables and…

Computation and Language · Computer Science 2025-10-21 Ilya Koziev

Lemmatization is the process of grouping together the inflected forms of a word so they can be analysed as a single item, identified by the word's lemma, or dictionary form. In computational linguistics, lemmatisation is the algorithmic…

Computation and Language · Computer Science 2022-07-26 Michal Karwatowski , Marcin Pietron

Semantics, morphology and syntax are strongly interdependent. However, the majority of computational methods for semantic change detection use distributional word representations which encode mostly semantics. We investigate an alternative…

Computation and Language · Computer Science 2021-09-23 Mario Giulianelli , Andrey Kutuzov , Lidia Pivovarova

This paper presents a grammar and style checker demonstrator for Spanish and Greek native writers developed within the project GramCheck. Besides a brief grammar error typology for Spanish, a linguistically motivated approach to detection…

cmp-lg · Computer Science 2016-08-15 Flora Ramírez Bustamante , Fernando Sánchez León

In natural language processing, the entropy of a language is a measure of its unpredictability and complexity. The first study on this subject was conducted by Claude Shannon in 1951. By having participants predict the next character in a…

Computation and Language · Computer Science 2026-05-01 Anton Lavreniuk , Mykyta Mudryi , Markiian Chaklosh