English
Related papers

Related papers: HBert + BiasCorp -- Fighting Racism on the Web

200 papers

Language model (LM) pretraining can learn various knowledge from text corpora, helping downstream tasks. However, existing methods such as BERT model a single document, and do not capture dependencies or knowledge that span across…

Computation and Language · Computer Science 2022-03-31 Michihiro Yasunaga , Jure Leskovec , Percy Liang

Recent advancements in natural language generation has raised serious concerns. High-performance language models are widely used for language generation tasks because they are able to produce fluent and meaningful sentences. These models…

Computation and Language · Computer Science 2020-10-06 Saurabh Gupta , Huy H. Nguyen , Junichi Yamagishi , Isao Echizen

Modern urbanization is demanding smarter technologies to improve a variety of applications in intelligent transportation systems to relieve the increasing amount of vehicular traffic congestion and incidents. Existing incident detection…

Computation and Language · Computer Science 2020-04-30 Xiangpeng Wan , Hakim Ghazzai , Yehia Massoud

In recent years, the introduction of the Transformer models sparked a revolution in natural language processing (NLP). BERT was one of the first text encoders using only the attention mechanism without any recurrent parts to achieve…

Computation and Language · Computer Science 2022-07-01 Ilan Perez , Raphael Reinauer

Academic researchers and social media entities grappling with the identification of hate speech face significant challenges, primarily due to the vast scale of data and the dynamic nature of hate speech. Given the ethical and practical…

Computation and Language · Computer Science 2024-05-08 Dengyi Liu , Minghao Wang , Andrew G. Catlin

Past studies in Sarcasm Detection mostly make use of Twitter datasets collected using hashtag-based supervision but such datasets are noisy in terms of labels and language. Furthermore, many tweets are replies to other tweets, and detecting…

Computation and Language · Computer Science 2022-12-13 Rishabh Misra

Toxic comments in online platforms are an unavoidable social issue under the cloak of anonymity. Hate speech detection has been actively done for languages such as English, German, or Italian, where manually labeled corpus has been…

Computation and Language · Computer Science 2020-05-27 Jihyung Moon , Won Ik Cho , Junbum Lee

Time is an important aspect of documents and is used in a range of NLP and IR tasks. In this work, we investigate methods for incorporating temporal information during pre-training to further improve the performance on time-related tasks.…

Computation and Language · Computer Science 2023-04-28 Jiexin Wang , Adam Jatowt , Masatoshi Yoshikawa , Yi Cai

The proliferation of radical content on online platforms poses significant risks, including inciting violence and spreading extremist ideologies. Despite ongoing research, existing datasets and models often fail to address the complexities…

Computation and Language · Computer Science 2024-12-20 Arij Riabi , Virginie Mouilleron , Menel Mahamdi , Wissam Antoun , Djamé Seddah

Social media often serves as a breeding ground for various hateful and offensive content. Identifying such content on social media is crucial due to its impact on the race, gender, or religion in an unprejudiced society. However, while…

Computation and Language · Computer Science 2022-10-10 Mithun Das , Somnath Banerjee , Punyajoy Saha , Animesh Mukherjee

The rapid growth of social media in recent years has fed into some highly undesirable phenomena such as proliferation of abusive and offensive language on the Internet. Previous research suggests that such hateful content tends to come from…

Computation and Language · Computer Science 2019-02-19 Pushkar Mishra , Marco Del Tredici , Helen Yannakoudakis , Ekaterina Shutova

This paper presents UniBERT, a compact multilingual language model that uses an innovative training framework that integrates three components: masked language modeling, adversarial training, and knowledge distillation. Pre-trained on a…

Computation and Language · Computer Science 2025-09-03 Andrei-Marius Avram , Marian Lupaşcu , Dumitru-Clementin Cercel , Ionuţ Mironică , Ştefan Trăuşan-Matu

In this research, we investigate techniques to detect hate speech in movies. We introduce a new dataset collected from the subtitles of six movies, where each utterance is annotated either as hate, offensive or normal. We apply transfer…

Computation and Language · Computer Science 2021-08-25 Niklas von Boguszewski , Sana Moin , Anirban Bhowmick , Seid Muhie Yimam , Chris Biemann

Hate speech is increasingly prevalent online, and its negative outcomes include increased prejudice, extremism, and even offline hate crime. Automatic detection of online hate speech can help us to better understand these impacts. However,…

Computation and Language · Computer Science 2021-02-10 John D Gallacher

The popularity of social media has created problems such as hate speech and sexism. The identification and classification of sexism in social media are very relevant tasks, as they would allow building a healthier social environment.…

Computation and Language · Computer Science 2021-11-09 Angel Felipe Magnossão de Paula , Roberto Fray da Silva , Ipek Baris Schlicht

In the era of social media platforms, identifying the credibility of online content is crucial to combat misinformation. We present the CREDiBERT (CREDibility assessment using Bi-directional Encoder Representations from Transformers), a…

Computation and Language · Computer Science 2024-02-20 Arash Amini , Yigit Ege Bayiz , Ashwin Ram , Radu Marculescu , Ufuk Topcu

Having a quality annotated corpus is essential especially for applied research. Despite the recent focus of Web science community on researching about cyberbullying, the community dose not still have standard benchmarks. In this paper, we…

Computation and Language · Computer Science 2018-05-25 Mohammadreza Rezvan , Saeedeh Shekarpour , Lakshika Balasuriya , Krishnaprasad Thirunarayan , Valerie Shalin , Amit Sheth

The present paper is about the participation of our team "techno" on CERIST'22 shared tasks. We used an available dataset "task1.c" related to covid-19 pandemic. It comprises 4128 tweets for sentiment analysis task and 8661 tweets for fake…

Computation and Language · Computer Science 2023-04-04 Rabia Bounaama , Mohammed El Amine Abderrahim

We describe AraNet, a collection of deep learning Arabic social media processing tools. Namely, we exploit an extensive host of publicly available and novel social media datasets to train bidirectional encoders from transformer models…

Computation and Language · Computer Science 2020-04-14 Muhammad Abdul-Mageed , Chiyu Zhang , Azadeh Hashemi , El Moatez Billah Nagoudi

Building on current work on multilingual hate speech (e.g., Ousidhoum et al. (2019)) and hate speech reduction (e.g., Sap et al. (2020)), we present XTREMESPEECH, a new hate speech dataset containing 20,297 social media passages from…

Computation and Language · Computer Science 2022-03-23 Antonis Maronikolakis , Axel Wisiorek , Leah Nann , Haris Jabbar , Sahana Udupa , Hinrich Schuetze
‹ Prev 1 4 5 6 7 8 10 Next ›