中文
相关论文

相关论文: Multi-Platform Aggregated Dataset of Online Commun…

200 篇论文

Online web communities often face bans for violating platform policies, encouraging their migration to alternative platforms. This migration, however, can result in increased toxicity and unforeseen consequences on the new platform. In…

社会与信息网络 · 计算机科学 2024-05-17 Jay Patel , Pujan Paudel , Emiliano De Cristofaro , Gianluca Stringhini , Jeremy Blackburn

Voat.co was a news aggregator website that shut down on December 25, 2020. The site had a troubled history and was known for hosting various banned subreddits. This paper presents a dataset with over 2.3M submissions and 16.2M comments…

社会与信息网络 · 计算机科学 2022-04-25 Amin Mekacher , Antonis Papasavva

The World Wide Web is a complex interconnected digital ecosystem, where information and attention flow between platforms and communities throughout the globe. These interactions co-construct how we understand the world, reflecting and…

计算机与社会 · 计算机科学 2025-04-17 Patrick Gildersleve , Anna Beers , Viviane Ito , Agustin Orozco , Francesca Tripodi

Proactively identifying misinformation spreaders is an important step towards mitigating the impact of fake news on our society. In this paper, we introduce a new contemporary Reddit dataset for fake news spreader analysis, called FACTOID,…

社会与信息网络 · 计算机科学 2022-05-13 Flora Sakketou , Joan Plepi , Riccardo Cervero , Henri-Jacques Geiss , Paolo Rosso , Lucie Flek

The Public Health Advocacy Dataset (PHAD) is a comprehensive collection of 5,730 videos related to tobacco products sourced from social media platforms like TikTok and YouTube. This dataset encompasses 4.3 million frames and includes…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Naga VS Raviteja Chappa , Charlotte McCormick , Susana Rodriguez Gongora , Page Daniel Dobbs , Khoa Luu

Pollution of online social spaces caused by rampaging d/misinformation is a growing societal concern. However, recent decisions to reduce access to social media APIs are causing a shortage of publicly available, recent, social media data,…

社会与信息网络 · 计算机科学 2024-11-25 Andrea Failla , Giulio Rossetti

Making online social communities 'better' is a challenging undertaking, as online communities are extraordinarily varied in their size, topical focus, and governance. As such, what is valued by one community may not be valued by another.…

社会与信息网络 · 计算机科学 2022-05-11 Galen Weld , Amy X. Zhang , Tim Althoff

Online hate speech is a recent problem in our society that is rising at a steady pace by leveraging the vulnerabilities of the corresponding regimes that characterise most social media platforms. This phenomenon is primarily fostered by…

计算与语言 · 计算机科学 2022-01-05 Ioannis Mollas , Zoe Chrysopoulou , Stamatis Karlos , Grigorios Tsoumakas

Understanding the sociodemographic composition of online platforms is essential for accurately interpreting digital behavior and its societal implications. Yet, current methods often lack the transparency and reliability required, risking…

社会与信息网络 · 计算机科学 2025-11-04 Federico Cinus , Corrado Monti , Paolo Bajardi , Gianmarco De Francisci Morales

In this dataset paper, we present a three-stage process to collect Reddit comments that are removed comments by moderators of several subreddits, for violating subreddit rules and guidelines. Other than the fact that these comments were…

社会与信息网络 · 计算机科学 2019-07-18 Eshwar Chandrasekharan , Eric Gilbert

A major task for moderators of online spaces is norm-setting, essentially creating shared norms for user behavior in their communities. Platform design principles emphasize the importance of highlighting norm-adhering examples and…

人机交互 · 计算机科学 2026-04-07 Agam Goyal , Charlotte Lambert , Yoshee Jain , Eshwar Chandrasekharan

This paper describes the development of a multilingual, manually annotated dataset for three under-resourced Dravidian languages generated from social media comments. The dataset was annotated for sentiment analysis and offensive language…

Toxic sentiment analysis on Twitter (X) often focuses on specific topics and events such as politics and elections. Datasets of toxic users in such research are typically gathered through lexicon-based techniques, providing only a…

社会与信息网络 · 计算机科学 2024-06-06 Hina Qayyum , Muhammad Ikram , Benjamin Zhao , Ian Wood , Mohamad Ali Kaafar , Nicolas Kourtellis

Social media platforms have become a hub for political activities and discussions, democratizing participation in these endeavors. However, they have also become an incubator for manipulation campaigns, like information operations (IOs).…

计算机与社会 · 计算机科学 2024-11-21 Ozgur Can Seckin , Manita Pote , Alexander Nwala , Lake Yin , Luca Luceri , Alessandro Flammini , Filippo Menczer

In this work, we construct and release a multi-domain and multi-modality event dataset (MMED), containing 25,165 textual news articles collected from hundreds of news media sites (e.g., Yahoo News, Google News, CNN News.) and 76,516 image…

多媒体 · 计算机科学 2019-04-10 Zhenguo Yang , Zehang Lin , Min Cheng , Qing Li , Wenyin Liu

This work presents MAD (Multimodal Affection Dataset), a multimodal emotion dataset designed for affective computing and neurophysiological modeling. MAD is built upon synchronous collection of diverse physiological signals (EEG, ECG, EOG,…

信号处理 · 电气工程与系统科学 2026-03-09 Shengwei Guo , Yunqing Qiao , Wenzhan Zhang , Bo Liu , Yong Wang , Guobing Sun

Since the Internet is flooded with hate, it is one of the main tasks for NLP experts to master automated online content moderation. However, advancements in this field require improved access to publicly available accurate and non-synthetic…

Social media are shifting towards pluralism -- community-governed platforms where groups define their own norms. What violates rules in one community may be perfectly acceptable in another. Can AI models help moderate such pluralistic…

计算与语言 · 计算机科学 2026-05-19 Zoher Kachwala , Bao Tran Truong , Rasika Muralidharan , Haewoon Kwak , Jisun An , Filippo Menczer

In this paper we present a benchmark dataset generated as part of a project for automatic identification of misogyny within online content, which focuses in particular on memes. The benchmark here described is composed of 800 memes…

人工智能 · 计算机科学 2022-10-07 Francesca Gasparini , Giulia Rizzi , Aurora Saibene , Elisabetta Fersini

We present a shared data model for enabling data science in Massive Open Online Courses (MOOCs). The model captures students interactions with the online platform. The data model is platform agnostic and is based on some basic core actions…

信息检索 · 计算机科学 2014-06-10 Kalyan Veeramachaneni , Sherif Halawa , Franck Dernoncourt , Una-May O'Reilly , Colin Taylor , Chuong Do
‹ 上一页 1 2 3 10 下一页 ›