中文
相关论文

相关论文: The Grammar and Syntax Based Corpus Analysis Tool …

200 篇论文

Large language models (LLMs) are known to exhibit biases in downstream tasks, especially when dealing with sensitive topics such as political discourse, gender identity, ethnic relations, or national stereotypes. Although significant…

计算与语言 · 计算机科学 2025-08-18 Martin Pavlíček , Tomáš Filip , Petr Sosík

Language students are most engaged while reading texts at an appropriate difficulty level. However, existing methods of evaluating text difficulty focus mainly on vocabulary and do not prioritize grammatical features, hence they do not work…

计算与语言 · 计算机科学 2017-02-17 Shuhan Wang , Erik Andersen

In this article, we describe some discursive segmentation methods as well as a preliminary evaluation of the segmentation quality. Although our experiment were carried for documents in French, we have developed three discursive segmentation…

Psychometric measures of ability, attitudes, perceptions, and beliefs are crucial for understanding user behaviors in various contexts including health, security, e-commerce, and finance. Traditionally, psychometric dimensions have been…

计算与语言 · 计算机科学 2020-07-28 Ahmed Abbasi , David G. Dobolyi , Richard G. Netemeyer

This paper presents the first Swedish evaluation benchmark for textual semantic similarity. The benchmark is compiled by simply running the English STS-B dataset through the Google machine translation API. This paper discusses potential…

计算与语言 · 计算机科学 2020-12-01 Tim Isbister , Magnus Sahlgren

Keyphrase extraction for morphologically rich, low-resource languages remains understudied, largely due to the scarcity of suitable evaluation datasets. We address this gap for Slovak by constructing a dataset of 227,432 scientific…

计算与语言 · 计算机科学 2026-03-17 David Števaňák , Marek Šuppa

Methods for learning sentence representations have been actively developed in recent years. However, the lack of pre-trained models and datasets annotated at the sentence level has been a problem for low-resource languages such as Polish…

计算与语言 · 计算机科学 2020-01-24 Sławomir Dadas , Michał Perełkiewicz , Rafał Poświata

Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based on Czech historical documents, containing human-defined…

计算与语言 · 计算机科学 2026-03-05 Martin Kostelník , Michal Hradiš , Martin Dočekal

This paper describes the speech processing activities conducted at the Polish consortium of the CLARIN project. The purpose of this segment of the project was to develop specific tools that would allow for automatic and semi-automatic…

计算与语言 · 计算机科学 2017-06-02 Danijel Koržinek , Krzysztof Marasek , Łukasz Brocki , Krzysztof Wołk

This paper presents Summary Workbench, a new tool for developing and evaluating text summarization models. New models and evaluation measures can be easily integrated as Docker-based plugins, allowing to examine the quality of their…

计算与语言 · 计算机科学 2022-10-19 Shahbaz Syed , Dominik Schwabe , Martin Potthast

Despite the extensive amount of labeled datasets in the NLP text classification field, the persistent imbalance in data availability across various languages remains evident. To support further fair development of NLP models, exploring the…

计算与语言 · 计算机科学 2025-02-06 Daryna Dementieva , Valeriia Khylenko , Georg Groh

Over the recent years, there has been a growing interest in developing new research evaluation methods that could go beyond the traditional citation-based metrics. This interest is motivated on one side by the wider availability or even…

数字图书馆 · 计算机科学 2016-11-17 Drahomira Herrmannova , Petr Knoth

This new research explores the effects of various training methods on a Polish to English Statistical Machine Translation system for medical texts. Various elements of the EMEA parallel text corpora from the OPUS project were used as the…

计算与语言 · 计算机科学 2015-09-30 Krzysztof Wołk , Krzysztof Marasek

The lack of automatic evaluation metrics tailored for SignWriting presents a significant obstacle in developing effective transcription and translation models for signed languages. This paper introduces a comprehensive suite of evaluation…

计算与语言 · 计算机科学 2024-10-18 Amit Moryossef , Rotem Zilberman , Ohad Langer

Text mining analysis of tweets gathered during Polish presidential election on May 10th, 2015. The project included implementation of engine to retrieve information from Twitter, building document corpora, corpora cleaning, and creating…

计算与语言 · 计算机科学 2019-11-05 Karol Chlasta

Generative poetry systems require effective tools for data engineering and automatic evaluation, particularly to assess how well a poem adheres to versification rules, such as the correct alternation of stressed and unstressed syllables and…

计算与语言 · 计算机科学 2025-10-21 Ilya Koziev

Lemmatization is the process of grouping together the inflected forms of a word so they can be analysed as a single item, identified by the word's lemma, or dictionary form. In computational linguistics, lemmatisation is the algorithmic…

计算与语言 · 计算机科学 2022-07-26 Michal Karwatowski , Marcin Pietron

Semantics, morphology and syntax are strongly interdependent. However, the majority of computational methods for semantic change detection use distributional word representations which encode mostly semantics. We investigate an alternative…

计算与语言 · 计算机科学 2021-09-23 Mario Giulianelli , Andrey Kutuzov , Lidia Pivovarova

This paper presents a grammar and style checker demonstrator for Spanish and Greek native writers developed within the project GramCheck. Besides a brief grammar error typology for Spanish, a linguistically motivated approach to detection…

cmp-lg · 计算机科学 2016-08-15 Flora Ramírez Bustamante , Fernando Sánchez León

In natural language processing, the entropy of a language is a measure of its unpredictability and complexity. The first study on this subject was conducted by Claude Shannon in 1951. By having participants predict the next character in a…

计算与语言 · 计算机科学 2026-05-01 Anton Lavreniuk , Mykyta Mudryi , Markiian Chaklosh