中文
相关论文

相关论文: CroSentiNews 2.0: A Sentence-Level News Sentiment …

200 篇论文

In this article, we present the first in depth linguistic study of human feelings. While there has been substantial research on incorporating some affective categories into linguistic analysis (e.g. sentiment, and to a lesser extent,…

计算与语言 · 计算机科学 2018-11-07 Advaith Siddharthan , Nicolas Cherbuin , Paul J. Eslinger , Kasia Kozlowska , Nora A. Murphy , Leroy Lowe

We present the development of a dataset for Kazakh named entity recognition. The dataset was built as there is a clear need for publicly available annotated corpora in Kazakh, as well as annotation guidelines containing straightforward--but…

计算与语言 · 计算机科学 2022-04-08 Rustem Yeshpanov , Yerbolat Khassanov , Huseyin Atakan Varol

In this paper, we present TwiSent, a sentiment analysis system for Twitter. Based on the topic searched, TwiSent collects tweets pertaining to it and categorizes them into the different polarity classes positive, negative and objective.…

信息检索 · 计算机科学 2012-09-19 Subhabrata Mukherjee , Akshat Malu , A. R. Balamurali , Pushpak Bhattacharyya

Currently, there are more than a dozen Russian-language corpora for sentiment analysis, differing in the source of the texts, domain, size, number and ratio of sentiment classes, and annotation method. This work examines publicly available…

计算与语言 · 计算机科学 2021-06-29 Evgeny Kotelnikov

Extracting who says what to whom is a crucial part in analyzing human communication in today's abundance of data such as online news articles. Yet, the lack of annotated data for this task in German news articles severely limits the quality…

计算与语言 · 计算机科学 2024-04-26 Fynn Petersen-Frey , Chris Biemann

Sentiment analysis for the Bengali language has attracted increasing research interest in recent years. However, progress remains constrained by the scarcity of large-scale and diverse annotated datasets. Although several Bengali sentiment…

计算与语言 · 计算机科学 2026-01-29 Akif Islam , Sujan Kumar Roy , Md. Ekramul Hamid

I present a tool which tells the quality of document or its usefulness based on annotations. Annotation may include comments, notes, observation, highlights, underline, explanation, question or help etc. comments are used for evaluative…

信息检索 · 计算机科学 2011-11-08 Archana Shukla

Sentiment Analysis (SA) is a major field of study in natural language processing, computational linguistics and information retrieval. Interest in SA has been constantly growing in both academia and industry over the recent years. Moreover,…

Automated event extraction in social science applications often requires corpus-level evaluations: for example, aggregating text predictions across metadata and unbiased estimates of recall. We combine corpus-level evaluation requirements…

计算与语言 · 计算机科学 2021-05-28 Andrew Halterman , Katherine A. Keith , Sheikh Muhammad Sarwar , Brendan O'Connor

Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10000 Arabic tweets and…

计算与语言 · 计算机科学 2025-05-26 Wajdi Zaghouani , Md. Rafiul Biswas

Citation sentimet analysis is one of the little studied tasks for scientometric analysis. For citation analysis, we developed eight datasets comprising citation sentences, which are manually annotated by us into three sentiment polarities…

计算与语言 · 计算机科学 2020-05-12 Vishal Vyas , Kumar Ravi , Vadlamani Ravi , V. Uma , Srirangaraj Setlur , Venu Govindaraju

The present paper is about the participation of our team "techno" on CERIST'22 shared tasks. We used an available dataset "task1.c" related to covid-19 pandemic. It comprises 4128 tweets for sentiment analysis task and 8661 tweets for fake…

计算与语言 · 计算机科学 2023-04-04 Rabia Bounaama , Mohammed El Amine Abderrahim

In sentiment analysis of longer texts, there may be a variety of topics discussed, of entities mentioned, and of sentiments expressed regarding each entity. We find a lack of studies exploring how such texts express their sentiment towards…

计算与语言 · 计算机科学 2024-09-18 Egil Rønningstad , Roman Klinger , Lilja Øvrelid , Erik Velldal

In this work we propose a novel annotation scheme which factors hate speech into five separate discursive categories. To evaluate our scheme, we construct a corpus of over 2.9M Twitter posts containing hateful expressions directed at Jews,…

计算与语言 · 计算机科学 2023-11-08 Gal Ron , Effi Levi , Odelia Oshri , Shaul R. Shenhav

We introduce a corpus of 7,032 sentences rated by human annotators for formality, informativeness, and implicature on a 1-7 scale. The corpus was annotated using Amazon Mechanical Turk. Reliability in the obtained judgments was examined by…

计算与语言 · 计算机科学 2016-09-29 Shibamouli Lahiri

Measuring how semantics of words change over time improves our understanding of how cultures and perspectives change. Diachronic word embeddings help us quantify this shift, although previous studies leveraged substantial temporally…

计算与语言 · 计算机科学 2025-06-17 David Dukić , Ana Barić , Marko Čuljak , Josip Jukić , Martin Tutek

We introduce XED, a multilingual fine-grained emotion dataset. The dataset consists of human-annotated Finnish (25k) and English sentences (30k), as well as projected annotations for 30 additional languages, providing new resources for many…

计算与语言 · 计算机科学 2020-11-09 Emily Öhman , Marc Pàmies , Kaisla Kajava , Jörg Tiedemann

In light of unprecedented increases in the popularity of the internet and social media, comment moderation has never been a more relevant task. Semi-automated comment moderation systems greatly aid human moderators by either automatically…

计算与语言 · 计算机科学 2022-11-14 Ravi Shekhar , Mladen Karan , Matthew Purver

Introduction: Microblogging websites have massed rich data sources for sentiment analysis and opinion mining. In this regard, sentiment classification has frequently proven inefficient because microblog posts typically lack syntactically…

计算与语言 · 计算机科学 2024-03-08 Mojtaba Mazoochi , Leila Rabiei , Farzaneh Rahmani , Zeinab Rajabi

This paper presents a collection of highly comparable web corpora of Slovenian, Croatian, Bosnian, Montenegrin, Serbian, Macedonian, and Bulgarian, covering thereby the whole spectrum of official languages in the South Slavic language…

计算与语言 · 计算机科学 2024-05-28 Nikola Ljubešić , Taja Kuzman