中文
相关论文

相关论文: HBert + BiasCorp -- Fighting Racism on the Web

200 篇论文

Language model (LM) pretraining can learn various knowledge from text corpora, helping downstream tasks. However, existing methods such as BERT model a single document, and do not capture dependencies or knowledge that span across…

计算与语言 · 计算机科学 2022-03-31 Michihiro Yasunaga , Jure Leskovec , Percy Liang

Recent advancements in natural language generation has raised serious concerns. High-performance language models are widely used for language generation tasks because they are able to produce fluent and meaningful sentences. These models…

计算与语言 · 计算机科学 2020-10-06 Saurabh Gupta , Huy H. Nguyen , Junichi Yamagishi , Isao Echizen

Modern urbanization is demanding smarter technologies to improve a variety of applications in intelligent transportation systems to relieve the increasing amount of vehicular traffic congestion and incidents. Existing incident detection…

计算与语言 · 计算机科学 2020-04-30 Xiangpeng Wan , Hakim Ghazzai , Yehia Massoud

In recent years, the introduction of the Transformer models sparked a revolution in natural language processing (NLP). BERT was one of the first text encoders using only the attention mechanism without any recurrent parts to achieve…

计算与语言 · 计算机科学 2022-07-01 Ilan Perez , Raphael Reinauer

Academic researchers and social media entities grappling with the identification of hate speech face significant challenges, primarily due to the vast scale of data and the dynamic nature of hate speech. Given the ethical and practical…

计算与语言 · 计算机科学 2024-05-08 Dengyi Liu , Minghao Wang , Andrew G. Catlin

Past studies in Sarcasm Detection mostly make use of Twitter datasets collected using hashtag-based supervision but such datasets are noisy in terms of labels and language. Furthermore, many tweets are replies to other tweets, and detecting…

计算与语言 · 计算机科学 2022-12-13 Rishabh Misra

Toxic comments in online platforms are an unavoidable social issue under the cloak of anonymity. Hate speech detection has been actively done for languages such as English, German, or Italian, where manually labeled corpus has been…

计算与语言 · 计算机科学 2020-05-27 Jihyung Moon , Won Ik Cho , Junbum Lee

Time is an important aspect of documents and is used in a range of NLP and IR tasks. In this work, we investigate methods for incorporating temporal information during pre-training to further improve the performance on time-related tasks.…

计算与语言 · 计算机科学 2023-04-28 Jiexin Wang , Adam Jatowt , Masatoshi Yoshikawa , Yi Cai

The proliferation of radical content on online platforms poses significant risks, including inciting violence and spreading extremist ideologies. Despite ongoing research, existing datasets and models often fail to address the complexities…

计算与语言 · 计算机科学 2024-12-20 Arij Riabi , Virginie Mouilleron , Menel Mahamdi , Wissam Antoun , Djamé Seddah

Social media often serves as a breeding ground for various hateful and offensive content. Identifying such content on social media is crucial due to its impact on the race, gender, or religion in an unprejudiced society. However, while…

计算与语言 · 计算机科学 2022-10-10 Mithun Das , Somnath Banerjee , Punyajoy Saha , Animesh Mukherjee

The rapid growth of social media in recent years has fed into some highly undesirable phenomena such as proliferation of abusive and offensive language on the Internet. Previous research suggests that such hateful content tends to come from…

计算与语言 · 计算机科学 2019-02-19 Pushkar Mishra , Marco Del Tredici , Helen Yannakoudakis , Ekaterina Shutova

This paper presents UniBERT, a compact multilingual language model that uses an innovative training framework that integrates three components: masked language modeling, adversarial training, and knowledge distillation. Pre-trained on a…

In this research, we investigate techniques to detect hate speech in movies. We introduce a new dataset collected from the subtitles of six movies, where each utterance is annotated either as hate, offensive or normal. We apply transfer…

计算与语言 · 计算机科学 2021-08-25 Niklas von Boguszewski , Sana Moin , Anirban Bhowmick , Seid Muhie Yimam , Chris Biemann

Hate speech is increasingly prevalent online, and its negative outcomes include increased prejudice, extremism, and even offline hate crime. Automatic detection of online hate speech can help us to better understand these impacts. However,…

计算与语言 · 计算机科学 2021-02-10 John D Gallacher

The popularity of social media has created problems such as hate speech and sexism. The identification and classification of sexism in social media are very relevant tasks, as they would allow building a healthier social environment.…

计算与语言 · 计算机科学 2021-11-09 Angel Felipe Magnossão de Paula , Roberto Fray da Silva , Ipek Baris Schlicht

In the era of social media platforms, identifying the credibility of online content is crucial to combat misinformation. We present the CREDiBERT (CREDibility assessment using Bi-directional Encoder Representations from Transformers), a…

计算与语言 · 计算机科学 2024-02-20 Arash Amini , Yigit Ege Bayiz , Ashwin Ram , Radu Marculescu , Ufuk Topcu

Having a quality annotated corpus is essential especially for applied research. Despite the recent focus of Web science community on researching about cyberbullying, the community dose not still have standard benchmarks. In this paper, we…

The present paper is about the participation of our team "techno" on CERIST'22 shared tasks. We used an available dataset "task1.c" related to covid-19 pandemic. It comprises 4128 tweets for sentiment analysis task and 8661 tweets for fake…

计算与语言 · 计算机科学 2023-04-04 Rabia Bounaama , Mohammed El Amine Abderrahim

We describe AraNet, a collection of deep learning Arabic social media processing tools. Namely, we exploit an extensive host of publicly available and novel social media datasets to train bidirectional encoders from transformer models…

计算与语言 · 计算机科学 2020-04-14 Muhammad Abdul-Mageed , Chiyu Zhang , Azadeh Hashemi , El Moatez Billah Nagoudi

Building on current work on multilingual hate speech (e.g., Ousidhoum et al. (2019)) and hate speech reduction (e.g., Sap et al. (2020)), we present XTREMESPEECH, a new hate speech dataset containing 20,297 social media passages from…

计算与语言 · 计算机科学 2022-03-23 Antonis Maronikolakis , Axel Wisiorek , Leah Nann , Haris Jabbar , Sahana Udupa , Hinrich Schuetze