English
Related papers

Related papers: Semi-automatic Generation of Multilingual Datasets…

200 papers

With social media communities increasingly becoming places where suicidal individuals post and congregate, natural language processing presents an exciting avenue for the development of automated suicide risk assessment systems. However,…

Computation and Language · Computer Science 2024-12-17 Max Lovitt , Haotian Ma , Song Wang , Yifan Peng

The problem of predicting the location of users on large social networks like Twitter has emerged from real-life applications such as social unrest detection and online marketing. Twitter user geolocation is a difficult and active research…

Machine Learning · Computer Science 2017-12-22 Tien Huu Do , Duc Minh Nguyen , Evaggelia Tsiligianni , Bruno Cornelis , Nikos Deligiannis

An identity denotes the role an individual or a group plays in highly differentiated contemporary societies. In this paper, our goal is to classify Twitter users based on their role identities. We first collect a coarse-grained public…

Social and Information Networks · Computer Science 2020-03-05 Binxuan Huang , Kathleen M. Carley

As open-ended human-chatbot interaction becomes commonplace, sensitive content detection gains importance. In this work, we propose a two stage semi-supervised approach to bootstrap large-scale data for automatic sensitive language…

Computation and Language · Computer Science 2018-12-03 Chandra Khatri , Behnam Hedayatnia , Rahul Goel , Anushree Venkatesh , Raefer Gabriel , Arindam Mandal

Automated stance detection and related machine learning methods can provide useful insights for media monitoring and academic research. Many of these approaches require annotated training datasets, which limits their applicability for…

Computation and Language · Computer Science 2026-03-19 Mark Mets , Andres Karjus , Indrek Ibrus , Maximilian Schich

Building models to detect vaccine attitudes on social media is challenging because of the composite, often intricate aspects involved, and the limited availability of annotated data. Existing approaches have relied heavily on supervised…

Computation and Language · Computer Science 2022-06-22 Lixing Zhu , Zheng Fang , Gabriele Pergola , Rob Procter , Yulan He

Semantic sentence embeddings are usually supervisedly built minimizing distances between pairs of embeddings of sentences labelled as semantically similar by annotators. Since big labelled datasets are rare, in particular for non-English…

Computation and Language · Computer Science 2021-10-06 Marco Di Giovanni , Marco Brambilla

Argument mining and stance detection are central to understanding how opinions are formed and contested in online discourse. However, most publicly available resources focus on mainstream platforms such as Twitter and Reddit, leaving…

Computation and Language · Computer Science 2026-02-17 Fathima Ameen , Danielle Brown , Manusha Malgareddy , Amanul Haque

Supervised learning algorithms are heavily reliant on annotated datasets to train machine learning models. However, the curation of the annotated datasets is laborious and time consuming due to the manual effort involved and has become a…

Computation and Language · Computer Science 2022-09-27 Ramya Tekumalla , Juan M. Banda

In this paper, we discuss the development of a multilingual dataset annotated with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur. The context, here, is defined by the…

Computation and Language · Computer Science 2021-11-23 Ritesh Kumar , Enakshi Nandi , Laishram Niranjana Devi , Shyam Ratan , Siddharth Singh , Akash Bhagat , Yogesh Dawer

Detecting personal health mentions on social media is essential to complement existing health surveillance systems. However, annotating data for detecting health mentions at a large scale is a challenging task. This research employs a…

Computation and Language · Computer Science 2022-12-13 Olanrewaju Tahir Aduragba , Jialin Yu , Alexandra I. Cristea

Many under-resourced languages require high-quality datasets for specific tasks such as offensive language detection, disinformation, or misinformation identification. However, the intricacies of the content may have a detrimental effect on…

Computation and Language · Computer Science 2023-11-20 Stetsenko Daria

Stance detection aims to identify whether the author of a text is in favor of, against, or neutral to a given target. The main challenge of this task comes two-fold: few-shot learning resulting from the varying targets and the lack of…

Computation and Language · Computer Science 2022-06-28 Yan Jiang , Jinhua Gao , Huawei Shen , Xueqi Cheng

The first objective towards the effective use of microblogging services such as Twitter for situational awareness during the emerging disasters is discovery of the disaster-related postings. Given the wide range of possible disasters, using…

Computation and Language · Computer Science 2016-10-13 Shanshan Zhang , Slobodan Vucetic

Detecting disclosures of individuals' employment status on social media can provide valuable information to match job seekers with suitable vacancies, offer social protection, or measure labor market flows. However, identifying such…

Computation and Language · Computer Science 2022-06-14 Manuel Tonneau , Dhaval Adjodah , João Palotti , Nir Grinberg , Samuel Fraiberger

Natural language understanding (NLU) is a task that enables machines to understand human language. Some tasks, such as stance detection and sentiment analysis, are closely related to individual subjective perspectives, thus termed…

Computation and Language · Computer Science 2025-02-20 Yunpeng Xiao , Youpeng Zhao , Kai Shu

Anticipating audience reaction towards a certain piece of text is integral to several facets of society ranging from politics, research, and commercial industries. Sentiment analysis (SA) is a useful natural language processing (NLP)…

Artificial Intelligence · Computer Science 2022-08-23 Anna Nguyen , Antonio Longa , Massimiliano Luca , Joe Kaul , Gabriel Lopez

Recent progress in language model pre-training has led to important improvements in Named Entity Recognition (NER). Nonetheless, this progress has been mainly tested in well-formatted documents such as news, Wikipedia, or scientific…

Computation and Language · Computer Science 2022-11-16 Asahi Ushio , Leonardo Neves , Vitor Silva , Francesco Barbieri , Jose Camacho-Collados

Opinion prediction on Twitter is challenging due to the transient nature of tweet content and neighbourhood context. In this paper, we model users' tweet posting behaviour as a temporal point process to jointly predict the posting time and…

Social and Information Networks · Computer Science 2020-05-28 Lixing Zhu , Yulan He , Deyu Zhou

Recent advances in natural language processing (NLP) in online social media are evidently owed to large-scale datasets. However, labeling, storing, and processing a large number of textual data points, e.g., tweets, has remained…

Computation and Language · Computer Science 2022-02-02 Toktam A. Oghaz , Ivan Garibay