中文
相关论文

相关论文: 4chan & 8chan embeddings

200 篇论文

4chan is a popular online imageboard which has been widely studied due to an observed concentration of far-right, antisemitic, racist, misogynistic, and otherwise hateful material being posted to the site, as well as the emergence of…

数字图书馆 · 计算机科学 2023-07-10 Jack H. Culbert

Online hate speech can harmfully impact individuals and groups, specifically on non-moderated platforms such as 4chan where users can post anonymous content. This work focuses on analysing and measuring the prevalence of online hate on…

计算与语言 · 计算机科学 2025-04-02 Adrian Bermudez-Villalva , Maryam Mehrnezhad , Ehsan Toreini

Online toxic content has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. A significant amount of research has been focused on detecting or analyzing toxic content using machine-learning…

计算与语言 · 计算机科学 2025-09-19 Gautam Kishore Shahi , Tim A. Majchrzak

Polarization has increased substantially in political discourse, contributing to a widening partisan divide. In this paper, we analyzed large-scale, real-world language use in Reddit communities (294,476,146 comments) and in news outlets…

计算与语言 · 计算机科学 2023-10-17 Nakwon Rim , Marc G. Berman , Yuan Chang Leong

There is an ongoing debate about how to moderate toxic speech on social media and the impact of content moderation on online discourse. This paper proposes and validates a methodology for measuring the content-moderation-induced distortions…

社会与信息网络 · 计算机科学 2026-03-04 Mahyar Habibi , Dirk Hovy , Carlo Schwarz

This paper presents a dataset with over 3.3M threads and 134.5M posts from the Politically Incorrect board (/pol/) of the imageboard forum 4chan, posted over a period of almost 3.5 years (June 2016-November 2019). To the best of our…

计算机与社会 · 计算机科学 2020-04-02 Antonis Papasavva , Savvas Zannettou , Emiliano De Cristofaro , Gianluca Stringhini , Jeremy Blackburn

Large pre-trained language models are often trained on large volumes of internet data, some of which may contain toxic or abusive language. Consequently, language models encode toxic information, which makes the real-world usage of these…

计算与语言 · 计算机科学 2021-12-16 Andrew Wang , Mohit Sudhakar , Yangfeng Ji

In-game toxic language becomes the hot potato in the gaming industry and community. There have been several online game toxicity analysis frameworks and models proposed. However, it is still challenging to detect toxicity due to the nature…

计算与语言 · 计算机科学 2025-04-02 Yuanzhe Jia , Weixuan Wu , Feiqi Cao , Soyeon Caren Han

Toxic speech, also known as hate speech, is regarded as one of the crucial issues plaguing online social media today. Most recent work on toxic speech detection is constrained to the modality of text and written conversations with very…

计算与语言 · 计算机科学 2022-04-05 Sreyan Ghosh , Samden Lepcha , S Sakshi , Rajiv Ratn Shah , S. Umesh

The discussion-board site 4chan has been part of the Internet's dark underbelly since its inception, and recent political events have put it increasingly in the spotlight. In particular, /pol/, the "Politically Incorrect" board, has been a…

Word embeddings or distributed representations of words are being used in various applications like machine translation, sentiment analysis, topic identification etc. Quality of word embeddings and performance of their applications depends…

计算与语言 · 计算机科学 2020-03-09 Erion Çano , Maurizio Morisio

This paper presents a characterization of AI-generated images shared on 4chan, examining how this anonymous online community is (mis-)using generative image technologies. Through a methodical data collection process, we gathered 900 images…

计算机与社会 · 计算机科学 2025-06-18 Parth Gaba , Emiliano De Cristofaro

The phenomenal growth on the internet has helped in empowering individual's expressions, but the misuse of freedom of expression has also led to the increase of various cyber crimes and anti-social activities. Hate speech is one such issue…

计算与语言 · 计算机科学 2020-06-01 Prashant Kapil , Asif Ekbal , Dipankar Das

The Web has become the main source for news acquisition. At the same time, news discussion has become more social: users can post comments on news articles or discuss news articles on other platforms like Reddit. These features empower and…

社会与信息网络 · 计算机科学 2020-05-19 Savvas Zannettou , Mai ElSherief , Elizabeth Belding , Shirin Nilizadeh , Gianluca Stringhini

The proliferation of social media platforms has led to an increase in the spread of hate speech, particularly targeting vulnerable communities. Unfortunately, existing methods for automatically identifying and blocking toxic language rely…

计算与语言 · 计算机科学 2025-02-24 Shiza Ali , Jeremy Blackburn , Gianluca Stringhini

Toxic language remains an ongoing challenge on social media platforms, presenting significant issues for users and communities. This paper provides a cross-topic and cross-lingual analysis of toxicity in Reddit conversations. We collect 1.5…

计算与语言 · 计算机科学 2024-04-30 Wondimagegnhue Tsegaye Tufa , Ilia Markov , Piek Vossen

Tackling toxic behavior in digital communication continues to be a pressing concern for both academics and industry professionals. While significant research has explored toxicity on platforms like social networks and discussion boards,…

计算与语言 · 计算机科学 2025-09-01 Naquee Rizwan , Nayandeep Deb , Sarthak Roy , Vishwajeet Singh Solanki , Kiran Garimella , Animesh Mukherjee

Toxicity classification for voice heavily relies on the semantic content of speech. We propose a novel framework that utilizes cross-modal learning to integrate the semantic embedding of text into a multilabel speech toxicity classifier…

计算与语言 · 计算机科学 2024-11-19 Joseph Liu , Mahesh Kumar Nandwana , Janne Pylkkönen , Hannes Heikinheimo , Morgan McGuire

Human language is colored by a broad range of topics, but existing text analysis tools only focus on a small number of them. We present Empath, a tool that can generate and validate new lexical categories on demand from a small set of seed…

计算与语言 · 计算机科学 2016-02-24 Ethan Fast , Binbin Chen , Michael Bernstein

We present a neural-network based approach to classifying online hate speech in general, as well as racist and sexist speech in particular. Using pre-trained word embeddings and max/mean pooling from simple, fully-connected transformations…

计算与语言 · 计算机科学 2018-09-28 Rohan Kshirsagar , Tyus Cukuvac , Kathleen McKeown , Susan McGregor
‹ 上一页 1 2 3 10 下一页 ›