中文
相关论文

相关论文: Multilingual Abusiveness Identification on Code-Mi…

200 篇论文

Gang violence is a severe issue in major cities across the U.S. and recent studies [Patton et al. 2017] have found evidence of social media communications that can be linked to such violence in communities with high rates of exposure to…

Content moderation is the process of screening and monitoring user-generated content online. It plays a crucial role in stopping content resulting from unacceptable behaviors such as hate speech, harassment, violence against specific…

计算与语言 · 计算机科学 2023-01-02 Álvaro Huertas-García , Alejandro Martín , Javier Huertas Tato , David Camacho

Social media is daily creating massive multimedia content with paired image and text, presenting the pressing need to automate the vision and language understanding for various multimodal classification tasks. Compared to the commonly…

计算与语言 · 计算机科学 2023-03-28 Chunpu Xu , Jing Li

The rise of social media platforms has led to an increase in cyber-aggressive behavior, encompassing a broad spectrum of hostile behavior, including cyberbullying, online harassment, and the dissemination of offensive and hate speech. These…

计算与语言 · 计算机科学 2024-12-31 Swapnil Mane , Suman Kundu , Rajesh Sharma

In today's digital world, social media plays a significant role in facilitating communication and content sharing. However, the exponential rise in user-generated content has led to challenges in maintaining a respectful online environment.…

计算与语言 · 计算机科学 2024-03-05 Mohammad Dehghani

The detection of sensitive content in large datasets is crucial for ensuring that shared and analysed data is free from harmful material. However, current moderation tools, such as external APIs, suffer from limitations in customisation,…

计算与语言 · 计算机科学 2025-06-25 Dimosthenis Antypas , Indira Sen , Carla Perez-Almendros , Jose Camacho-Collados , Francesco Barbieri

Online communities have gained considerable importance in recent years due to the increasing number of people connected to the Internet. Moderating user content in online communities is mainly performed manually, and reducing the workload…

信息检索 · 计算机科学 2019-01-16 Etienne Papegnies , Vincent Labatut , Richard Dufour , Georges Linares

Harmful content detection models tend to have higher false positive rates for content from marginalized groups. In the context of marginal abuse modeling on Twitter, such disproportionate penalization poses the risk of reduced visibility,…

计算与语言 · 计算机科学 2022-10-13 Kyra Yee , Alice Schoenauer Sebag , Olivia Redfield , Emily Sheng , Matthias Eck , Luca Belli

Offensive language is pervasive in social media. Individuals frequently take advantage of the perceived anonymity of computer-mediated communication, using this to engage in behavior that many of them would not consider in real life. The…

计算与语言 · 计算机科学 2021-04-13 Nikhil Oswal

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

We address fine-grained multilingual language identification: providing a language code for every token in a sentence, including codemixed text containing multiple languages. Such text is prevalent online, in documents, social media, and…

计算与语言 · 计算机科学 2018-10-10 Yuan Zhang , Jason Riesa , Daniel Gillick , Anton Bakalov , Jason Baldridge , David Weiss

The automatic identification of harmful content online is of major concern for social media platforms, policymakers, and society. Researchers have studied textual, visual, and audio content, but typically in isolation. Yet, harmful content…

The increase in abusive content on online social media platforms is impacting the social life of online users. Use of offensive and hate speech has been making so-cial media toxic. Homophobia and transphobia constitute offensive comments…

计算与语言 · 计算机科学 2023-04-05 Deepawali Sharma , Vedika Gupta , Vivek Kumar Singh

Social media platforms have a vital role in the modern world, serving as conduits for communication, the exchange of ideas, and the establishment of networks. However, the misuse of these platforms through toxic comments, which can range…

计算与语言 · 计算机科学 2025-06-24 Mukaffi Bin Moin , Pronay Debnath , Usafa Akther Rifa , Rijeet Bin Anis

Online abusive content detection, particularly in low-resource settings and within the audio modality, remains underexplored. We investigate the potential of pre-trained audio representations for detecting abusive language in low-resource…

计算与语言 · 计算机科学 2024-12-16 Aditya Narayan Sankaran , Reza Farahbakhsh , Noel Crespi

Social media has penetrated into multilingual societies, however most of them use English to be a preferred language for communication. So it looks natural for them to mix their cultural language with English during conversations resulting…

计算与语言 · 计算机科学 2020-10-16 Shubhanker Banerjee , Arun Jayapal , Sajeetha Thavareesan

Social media have been deliberately used for malicious purposes, including political manipulation and disinformation. Most research focuses on high-resource languages. However, malicious actors share content across countries and languages,…

社会与信息网络 · 计算机科学 2023-05-04 Samar Haider , Luca Luceri , Ashok Deb , Adam Badawy , Nanyun Peng , Emilio Ferrara

In today's global digital landscape, misinformation transcends linguistic boundaries, posing a significant challenge for moderation systems. Most approaches to misinformation detection are monolingual, focused on high-resource languages,…

计算与语言 · 计算机科学 2025-04-01 Xinyu Wang , Wenbo Zhang , Sarah Rajtmajer

Language Identification (LID) is a core task in multilingual NLP, yet current systems often overfit to clean, monolingual data. This work introduces DIVERS-BENCH, a comprehensive evaluation of state-of-the-art LID models across diverse…

计算与语言 · 计算机科学 2025-09-23 Jessica Ojo , Zina Kamel , David Ifeoluwa Adelani

In order to study online hate speech, the availability of datasets containing the linguistic phenomena of interest are of crucial importance. However, when it comes to specific target groups, for example teenagers, collecting such data may…

计算与语言 · 计算机科学 2020-05-06 Alessio Palmero Aprosio , Stefano Menini , Sara Tonelli