English
Related papers

Related papers: Czech Dataset for Cross-lingual Subjectivity Class…

200 papers

This paper describes a novel dataset consisting of sentences with semantic similarity annotations. The data originate from the journalistic domain in the Czech language. We describe the process of collecting and annotating the data in…

Computation and Language · Computer Science 2022-01-24 Jakub Sido , Michal Seják , Ondřej Pražák , Miloslav Konopík , Václav Moravec

In this paper, we introduce a novel Czech dataset for aspect-based sentiment analysis (ABSA), which consists of 3.1K manually annotated reviews from the restaurant domain. The dataset is built upon the older Czech dataset, which contained…

Computation and Language · Computer Science 2025-08-12 Jakub Šmíd , Pavel Přibáň , Ondřej Pražák , Pavel Král

In this paper, we aim at improving Czech sentiment with transformer-based models and their multilingual versions. More concretely, we study the task of polarity detection for the Czech language on three sentiment polarity datasets. We…

Computation and Language · Computer Science 2021-08-25 Pavel Přibáň , Josef Steinberger

This paper describes the training process of the first Czech monolingual language representation models based on BERT and ALBERT architectures. We pre-train our models on more than 340K of sentences, which is 50 times more than multilingual…

Computation and Language · Computer Science 2021-08-23 Jakub Sido , Ondřej Pražák , Pavel Přibáň , Jan Pašek , Michal Seják , Miloslav Konopík

This paper introduces a novel Czech dataset in the restaurant domain for aspect-based sentiment analysis (ABSA), enriched with annotations of opinion terms. The dataset supports three distinct ABSA tasks involving opinion terms,…

Computation and Language · Computer Science 2026-03-05 Jakub Šmíd , Pavel Přibáň , Pavel Král

Reading comprehension is a well studied task, with huge training datasets in English. This work focuses on building reading comprehension systems for Czech, without requiring any manually annotated Czech training data. First of all, we…

Computation and Language · Computer Science 2020-07-06 Kateřina Macková , Milan Straka

This paper deals with cross-lingual sentiment analysis in Czech, English and French languages. We perform zero-shot cross-lingual classification using five linear transformations combined with LSTM and CNN based classifiers. We compare the…

Computation and Language · Computer Science 2022-09-16 Pavel Přibáň , Jakub Šmíd , Adam Mištera , Pavel Král

This article presents the creation of an Estonian-language dataset for document-level subjectivity, analyzes the resulting annotations, and reports an initial experiment of automatic subjectivity analysis using a large language model (LLM).…

Computation and Language · Computer Science 2025-12-11 Karl Gustav Gailit , Kadri Muischnek , Kairit Sirts

We introduce XED, a multilingual fine-grained emotion dataset. The dataset consists of human-annotated Finnish (25k) and English sentences (30k), as well as projected annotations for 30 additional languages, providing new resources for many…

Computation and Language · Computer Science 2020-11-09 Emily Öhman , Marc Pàmies , Kaisla Kajava , Jörg Tiedemann

Pre-trained models for Czech Natural Language Processing are often evaluated on purely linguistic tasks (POS tagging, parsing, NER) and relatively simple classification tasks such as sentiment classification or article classification from a…

Computation and Language · Computer Science 2023-07-21 Hynek Kydlíček , Jindřich Libovický

We present a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in a target language. The corpus covers twelve language pairs and directions for four European languages,…

Computation and Language · Computer Science 2022-02-22 Laura Perez-Beltrachini , Mirella Lapata

We introduce a new dataset for multi-class emotion analysis from long-form narratives in English. The Dataset for Emotions of Narrative Sequences (DENS) was collected from both classic literature available on Project Gutenberg and modern…

Computation and Language · Computer Science 2019-10-28 Chen Liu , Muhammad Osama , Anderson de Andrade

This study addresses a binary classification task to determine whether a text sequence, either a sentence or paragraph, is subjective or objective. The task spans five languages: Arabic, Bulgarian, English, German, and Italian, along with a…

Computation and Language · Computer Science 2024-07-16 Md. Rafiul Biswas , Abrar Tasneem Abir , Wajdi Zaghouani

Emotion detection can provide us with a window into understanding human behavior. Due to the complex dynamics of human emotions, however, constructing annotated datasets to train automated models can be expensive. Thus, we explore the…

Computation and Language · Computer Science 2022-05-06 Sabit Hassan , Shaden Shaar , Kareem Darwish

In this paper, we compare Czech-specific and multilingual sentence embedding models through intrinsic and extrinsic evaluation paradigms. For intrinsic evaluation, we employ Costra, a complex sentence transformation dataset, and several…

Computation and Language · Computer Science 2025-06-26 Petra Barančíková , Ondřej Bojar

Scaling semantic parsing models for task-oriented dialog systems to new languages is often expensive and time-consuming due to the lack of available datasets. Available datasets suffer from several shortcomings: a) they contain few…

Computation and Language · Computer Science 2021-01-28 Haoran Li , Abhinav Arora , Shuohui Chen , Anchit Gupta , Sonal Gupta , Yashar Mehdad

Topic localization aims to identify spans of text that express a given topic defined by a name and description. To study this task, we introduce a human-annotated benchmark based on Czech historical documents, containing human-defined…

Computation and Language · Computer Science 2026-03-05 Martin Kostelník , Michal Hradiš , Martin Dočekal

We develop novel annotation guidelines for sentence-level subjectivity detection, which are not limited to language-specific cues. We use our guidelines to collect NewsSD-ENG, a corpus of 638 objective and 411 subjective sentences extracted…

Neural Machine Translation (NMT) has improved translation by using Transformer-based models, but it still struggles with word ambiguity and context. This problem is especially important in domain-specific applications, which often have…

Computation and Language · Computer Science 2025-06-10 Mikołaj Pokrywka , Wojciech Kusa , Mieszko Rutkowski , Mikołaj Koszowski

Relation classification is one of the key topics in information extraction, which can be used to construct knowledge bases or to provide useful information for question answering. Current approaches for relation classification are mainly…

Computation and Language · Computer Science 2020-10-20 Abdullatif Köksal , Arzucan Özgür
‹ Prev 1 2 3 10 Next ›