English
Related papers

Related papers: CroSentiNews 2.0: A Sentence-Level News Sentiment …

200 papers

Fine-grained financial sentiment analysis on news headlines is a challenging task requiring human-annotated datasets to achieve high performance. Limited studies have tried to address the sentiment extraction task in a setting where…

Computation and Language · Computer Science 2023-05-23 Ankur Sinha , Satishwar Kedas , Rishu Kumar , Pekka Malo

Despite the importance of understanding causality, corpora addressing causal relations are limited. There is a discrepancy between existing annotation guidelines of event causality and conventional causality corpora that focus more on…

Having a quality annotated corpus is essential especially for applied research. Despite the recent focus of Web science community on researching about cyberbullying, the community dose not still have standard benchmarks. In this paper, we…

Computation and Language · Computer Science 2018-05-25 Mohammadreza Rezvan , Saeedeh Shekarpour , Lakshika Balasuriya , Krishnaprasad Thirunarayan , Valerie Shalin , Amit Sheth

Understanding how individuals perceive and react to information is fundamental for advancing social and behavioral sciences and developing human-centered AI systems. Current approaches often lack the granular data needed to model these…

Computation and Language · Computer Science 2025-07-08 Tiancheng Hu , Nigel Collier

Since state-of-the-art approaches to offensive language detection rely on supervised learning, it is crucial to quickly adapt them to the continuously evolving scenario of social media. While several approaches have been proposed to tackle…

Computation and Language · Computer Science 2022-10-17 Elisa Leonardelli , Stefano Menini , Alessio Palmero Aprosio , Marco Guerini , Sara Tonelli

Newsletters and social networks can reflect the opinion about the market and specific stocks from the perspective of analysts and the general public on products and/or services provided by a company. Therefore, sentiment analysis of these…

Computation and Language · Computer Science 2021-12-28 Elvys Linhares Pontes , Mohamed Benjannet

This paper presents the Norwegian Review Corpus (NoReC), created for training and evaluating models for document-level sentiment analysis. The full-text reviews have been collected from major Norwegian news sources and cover a range of…

Computation and Language · Computer Science 2017-10-17 Erik Velldal , Lilja Øvrelid , Eivind Alexander Bergem , Cathrine Stadsnes , Samia Touileb , Fredrik Jørgensen

This paper presents Hotter and Colder, a dataset designed to analyze various types of online behavior in Icelandic blog comments. Building on previous work, we used GPT-4o mini to annotate approximately 800,000 comments for 25 tasks,…

Computation and Language · Computer Science 2025-02-25 Steinunn Rut Friðriksdóttir , Dan Saattrup Nielsen , Hafsteinn Einarsson

This paper describes the development of a multilingual, manually annotated dataset for three under-resourced Dravidian languages generated from social media comments. The dataset was annotated for sentiment analysis and offensive language…

Controllable text simplification is a crucial assistive technique for language learning and teaching. One of the primary factors hindering its advancement is the lack of a corpus annotated with sentence difficulty levels based on language…

Computation and Language · Computer Science 2022-10-24 Yuki Arase , Satoru Uchida , Tomoyuki Kajiwara

In this paper, we present a dataset containing 9,973 tweets related to the MeToo movement that were manually annotated for five different linguistic aspects: relevance, stance, hate speech, sarcasm, and dialogue acts. We present a detailed…

Computation and Language · Computer Science 2020-04-21 Akash Gautam , Puneet Mathur , Rakesh Gosangi , Debanjan Mahata , Ramit Sawhney , Rajiv Ratn Shah

Current TSA evaluation in a cross-domain setup is restricted to the small set of review domains available in existing datasets. Such an evaluation is limited, and may not reflect true performance on sites like Amazon or Yelp that host…

Computation and Language · Computer Science 2021-09-14 Matan Orbach , Orith Toledo-Ronen , Artem Spector , Ranit Aharonov , Yoav Katz , Noam Slonim

We introduce a new dataset for multi-class emotion analysis from long-form narratives in English. The Dataset for Emotions of Narrative Sequences (DENS) was collected from both classic literature available on Project Gutenberg and modern…

Computation and Language · Computer Science 2019-10-28 Chen Liu , Muhammad Osama , Anderson de Andrade

We introduce FinLin, a novel corpus containing investor reports, company reports, news articles, and microblogs from StockTwits, targeting multiple entities stemming from the automobile industry and covering a 3-month period. FinLin was…

Computation and Language · Computer Science 2020-03-10 Tobias Daudert

Sentiment analysis that classifies data into positive or negative has been dominantly used to recognize emotional aspects of texts, despite the deficit of thorough examination of emotional meanings. Recently, corpora labeled with more than…

Computation and Language · Computer Science 2022-05-12 Duyoung Jeon , Junho Lee , Cheongtag Kim

The continuous and increasing use of social media has enabled the expression of human thoughts, opinions, and everyday actions publicly at an unprecedented scale. We present the Vent dataset, the largest annotated dataset of text, emotions,…

Social and Information Networks · Computer Science 2019-03-26 Nikolaos Lykousas , Costantinos Patsakis , Andreas Kaltenbrunner , Vicenç Gómez

Argumentative stance classification plays a key role in identifying authors' viewpoints on specific topics. However, generating diverse pairs of argumentative sentences across various domains is challenging. Existing benchmarks often come…

Computation and Language · Computer Science 2024-11-19 Jiaqing Yuan , Ruijie Xi , Munindar P. Singh

In this paper, we introduce a new WordNet based similarity metric, SenSim, which incorporates sentiment content (i.e., degree of positive or negative sentiment) of the words being compared to measure the similarity between them. The…

Information Retrieval · Computer Science 2012-09-19 A. R. Balamurali , Subhabrata Mukherjee , Akshat Malu , Pushpak Bhattacharyya

Extensive research on target-dependent sentiment classification (TSC) has led to strong classification performances in domains where authors tend to explicitly express sentiment about specific entities or topics, such as in reviews or on…

Computation and Language · Computer Science 2021-05-21 Felix Hamborg , Karsten Donnay , Bela Gipp

We describe a gold standard corpus of protest events that comprise of various local and international sources from various countries in English. The corpus contains document, sentence, and token level annotations. This corpus facilitates…

Computation and Language · Computer Science 2020-08-04 Ali Hürriyetoğlu , Erdem Yörük , Deniz Yüret , Osman Mutlu , Çağrı Yoltar , Fırat Duruşan , Burak Gürel