中文
相关论文

相关论文: Chinese Cyberbullying Detection: Dataset, Method, …

200 篇论文

Chinese sequence labeling tasks are heavily reliant on accurate word boundary demarcation. Although current pre-trained language models (PLMs) have achieved substantial gains on these tasks, they rarely explicitly incorporate boundary…

计算与语言 · 计算机科学 2024-04-09 Longhui Zhang , Dingkun Long , Meishan Zhang , Yanzhao Zhang , Pengjun Xie , Min Zhang

Social media sites such as YouTube and Facebook have become an integral part of everyone's life and in the last few years, hate speech in the social media comment section has increased rapidly. Detection of hate speech on social media…

计算与语言 · 计算机科学 2020-12-18 Nauros Romim , Mosahed Ahmed , Hriteshwar Talukder , Md Saiful Islam

Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are essential for advancing counterspeech research. However,…

Named Entity Recognition and Relation Extraction for Chinese literature text is regarded as the highly difficult problem, partially because of the lack of tagging sets. In this paper, we build a discourse-level dataset from hundreds of…

计算与语言 · 计算机科学 2019-06-12 Jingjing Xu , Ji Wen , Xu Sun , Qi Su

This study used text mining to investigate people's insights about cyberbullying. English-language tweets were collected and analyzed by R software. Our analysis demonstrated three major themes: (a) the major actions that needed to be taken…

计算机与社会 · 计算机科学 2019-07-25 Iman Tahamtan , Li-Min Huang

We present NewsBench, a novel evaluation framework to systematically assess the capabilities of Large Language Models (LLMs) for editorial capabilities in Chinese journalism. Our constructed benchmark dataset is focused on four facets of…

计算与语言 · 计算机科学 2024-06-05 Miao Li , Ming-Bin Chen , Bo Tang , Shengbin Hou , Pengyu Wang , Haiying Deng , Zhiyu Li , Feiyu Xiong , Keming Mao , Peng Cheng , Yi Luo

Detoxifying offensive language while preserving the speaker's original intent is a challenging yet critical goal for improving the quality of online interactions. Although large language models (LLMs) show promise in rewriting toxic…

计算与语言 · 计算机科学 2025-05-22 Xintong Wang , Yixiao Liu , Jingheng Pan , Liang Ding , Longyue Wang , Chris Biemann

The proliferation of hate speech and offensive comments on social media has become increasingly prevalent due to user activities. Such comments can have detrimental effects on individuals' psychological well-being and social behavior. While…

With the recent rise of #MeToo, an increasing number of personal stories about sexual harassment and sexual abuse have been shared online. In order to push forward the fight against such harassment and abuse, we present the task of…

计算与语言 · 计算机科学 2018-09-17 Sweta Karlekar , Mohit Bansal

Internet memes have gained significant influence in communicating political, psychological, and sociocultural ideas. While memes are often humorous, there has been a rise in the use of memes for trolling and cyberbullying. Although a wide…

计算与语言 · 计算机科学 2024-01-19 Prince Jha , Krishanu Maity , Raghav Jain , Apoorv Verma , Sriparna Saha , Pushpak Bhattacharyya

Internet censorship measurements rely on lists of websites to be tested, or "block lists" that are curated by third parties. Unfortunately, many of these lists are not public, and those that are tend to focus on a small group of topics,…

计算机与社会 · 计算机科学 2018-07-20 Austin Hounsel , Prateek Mittal , Nick Feamster

Large language models (LLMs), like ChatGPT and GPT-4, have demonstrated remarkable abilities in natural language understanding and generation. However, alongside their positive impact on our daily tasks, they can also produce harmful…

计算与语言 · 计算机科学 2023-10-10 Liang Xu , Kangkang Zhao , Lei Zhu , Hang Xue

Recent works of opinion expression identification (OEI) rely heavily on the quality and scale of the manually-constructed training corpus, which could be extremely difficult to satisfy. Crowdsourcing is one practical solution for this…

计算与语言 · 计算机科学 2022-04-25 Xin Zhang , Guangwei Xu , Yueheng Sun , Meishan Zhang , Xiaobin Wang , Min Zhang

The research described in this paper concerns automatic cyberbullying detection in social media. There are two goals to achieve: building a gold standard cyberbullying detection dataset and measuring the performance of the Samurai…

计算与语言 · 计算机科学 2018-08-03 Michał Ptaszyński , Gniewosz Leliwa , Mateusz Piech , Aleksander Smywiński-Pohl

Cyberbullying, identified as intended and repeated online bullying behavior, has become increasingly prevalent in the past few decades. Despite the significant progress made thus far, the focus of most existing work on cyberbullying…

计算与语言 · 计算机科学 2021-02-12 Suyu Ge , Lu Cheng , Huan Liu

We present CFEVER, a Chinese dataset designed for Fact Extraction and VERification. CFEVER comprises 30,012 manually created claims based on content in Chinese Wikipedia. Each claim in CFEVER is labeled as "Supports", "Refutes", or "Not…

计算与语言 · 计算机科学 2025-06-17 Ying-Jia Lin , Chun-Yi Lin , Chia-Jen Yeh , Yi-Ting Li , Yun-Yu Hu , Chih-Hao Hsu , Mei-Feng Lee , Hung-Yu Kao

Many studies have demonstrated that large language models (LLMs) can produce harmful responses, exposing users to unexpected risks when LLMs are deployed. Previous studies have proposed comprehensive taxonomies of the risks posed by LLMs,…

计算与语言 · 计算机科学 2024-08-06 Yuxia Wang , Zenan Zhai , Haonan Li , Xudong Han , Lizhi Lin , Zhenxuan Zhang , Jingru Zhao , Preslav Nakov , Timothy Baldwin

Dialogue agents have been receiving increasing attention for years, and this trend has been further boosted by the recent progress of large language models (LLMs). Stance detection and dialogue summarization are two core tasks of dialogue…

计算与语言 · 计算机科学 2024-10-18 Xiutian Zhao , Ke Wang , Wei Peng

Algorithms are widely applied to detect hate speech and abusive language in social media. We investigated whether the human-annotated data used to train these algorithms are biased. We utilized a publicly available annotated Twitter dataset…

计算与语言 · 计算机科学 2020-05-29 Jae Yeon Kim , Carlos Ortiz , Sarah Nam , Sarah Santiago , Vivek Datta

Detecting harmful content on social media, such as Twitter, is made difficult by the fact that the seemingly simple yes/no classification conceals a significant amount of complexity. Unfortunately, while several datasets have been collected…

计算与语言 · 计算机科学 2023-11-14 Saad Almohaimeed , Saleh Almohaimeed , Ashfaq Ali Shafin , Bogdan Carbunar , Ladislau Bölöni