中文
相关论文

相关论文: CoRAL: a Context-aware Croatian Abusive Language D…

200 篇论文

State-of-the-art conversational AI systems raise concerns due to their potential risks of generating unsafe, toxic, unethical, or dangerous content. Previous works have developed datasets to teach conversational agents the appropriate…

计算与语言 · 计算机科学 2024-02-02 Souvik Das , Rohini K. Srihari

Active learning (AL), which aims to construct an effective training set by iteratively curating the most formative unlabeled data for annotation, has been widely used in low-resource tasks. Most active learning techniques in classification…

计算与语言 · 计算机科学 2024-12-17 Yun Luo , Zhen Yang , Fandong Meng , Yingjie Li , Fang Guo , Qinglin Qi , Jie Zhou , Yue Zhang

Accurate detection and classification of online hate is a difficult task. Implicit hate is particularly challenging as such content tends to have unusual syntax, polysemic words, and fewer markers of prejudice (e.g., slurs). This problem is…

计算与语言 · 计算机科学 2021-06-11 Austin Botelho , Bertie Vidgen , Scott A. Hale

We present COPAL-ID, a novel, public Indonesian language common sense reasoning dataset. Unlike the previous Indonesian COPA dataset (XCOPA-ID), COPAL-ID incorporates Indonesian local and cultural nuances, and therefore, provides a more…

Social media platforms are being increasingly used by malicious actors to share unsafe content, such as images depicting sexual activity, cyberbullying, and self-harm. Consequently, major platforms use artificial intelligence (AI) and human…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Mazal Bethany , Brandon Wherry , Nishant Vishwamitra , Peyman Najafirad

The detection of sensitive content in large datasets is crucial for ensuring that shared and analysed data is free from harmful material. However, current moderation tools, such as external APIs, suffer from limitations in customisation,…

计算与语言 · 计算机科学 2025-06-25 Dimosthenis Antypas , Indira Sen , Carla Perez-Almendros , Jose Camacho-Collados , Francesco Barbieri

Multimodal learning seeks to integrate information from heterogeneous sources, where signals may be shared across modalities, specific to individual modalities, or emerge only through their interaction. While self-supervised multimodal…

机器学习 · 计算机科学 2026-02-17 Carolin Cissee , Raneen Younis , Zahra Ahmadi

Establishing whether language models can use contextual information in a human-plausible way is important to ensure their trustworthiness in real-world settings. However, the questions of when and which parts of the context affect model…

计算与语言 · 计算机科学 2024-03-14 Gabriele Sarti , Grzegorz Chrupała , Malvina Nissim , Arianna Bisazza

On-the-fly reasoning often requires adaptation to novel problems under limited data and distribution shift. This work introduces CausalARC: an experimental testbed for AI reasoning in low-data and out-of-distribution regimes, modeled after…

人工智能 · 计算机科学 2026-03-20 Jacqueline Maasch , John Kalantari , Kia Khezeli

Code-mixed discourse combines multiple languages in a single text. It is commonly used in informal discourse in countries with several official languages, but also in many other countries in combination with English or neighboring…

计算与语言 · 计算机科学 2025-04-16 Anjali Yadav , Tanya Garg , Matej Klemen , Matej Ulcar , Basant Agarwal , Marko Robnik Sikonja

Sarcasm is an intricate form of speech, where meaning is conveyed implicitly. Being a convoluted form of expression, detecting sarcasm is an assiduous problem. The difficulty in recognition of sarcasm has many pitfalls, including…

计算与语言 · 计算机科学 2020-06-02 Kartikey Pant , Tanvi Dadu

We study the impact of content moderation policies in online communities. In our theoretical model, a platform chooses a content moderation policy and individuals choose whether or not to participate in the community according to the…

数据结构与算法 · 计算机科学 2023-10-17 Cynthia Dwork , Chris Hays , Jon Kleinberg , Manish Raghavan

Recent advances in natural language processing have enabled the increasing use of text data in causal inference, particularly for adjusting confounding factors in treatment effect estimation. Although high-dimensional text can encode rich…

机器学习 · 计算机科学 2025-12-08 Lijinghua Zhang , Hengrui Cai

Moderation of user-generated content in an online community is a challenge that has great socio-economical ramifications. However, the costs incurred by delegating this work to human agents are high. For this reason, an automatic system…

信息检索 · 计算机科学 2019-02-01 Etienne Papegnies , Vincent Labatut , Richard Dufour , Georges Linares

The study of online discourse has become central to understanding societal polarization. While much research has focused on detecting overt toxicity, the subtle dynamics of social cohesion, meaning the interaction between divisive and…

计算与语言 · 计算机科学 2026-05-22 Aisha Ali Al-Athba , Wajdi Zaghouani

Whereas much of the success of the current generation of neural language models has been driven by increasingly large training corpora, relatively little research has been dedicated to analyzing these massive sources of textual data. In…

计算与语言 · 计算机科学 2021-06-02 Alexandra Sasha Luccioni , Joseph D. Viviano

Offensive content is pervasive in social media and a reason for concern to companies and government organizations. Several studies have been recently published investigating methods to detect the various forms of such content (e.g. hate…

计算与语言 · 计算机科学 2021-05-21 Tharindu Ranasinghe , Marcos Zampieri

Despite the valuable social interactions that online media promote, these systems provide space for speech that would be potentially detrimental to different groups of people. The moderation of content imposed by many social media has…

社会与信息网络 · 计算机科学 2021-08-30 Lucas Henrique Costa de Lima , Julio Reis , Philipe Melo , Fabricio Murai , Fabricio Benevenuto

Recent advancements in language model technology have significantly enhanced the ability to edit factual information. Yet, the modification of moral judgments, a crucial aspect of aligning models with human values, has garnered less…

人工智能 · 计算机科学 2026-03-31 Michael Ripa , Jim Davies

Automatic speech recognition (ASR) system is becoming a ubiquitous technology. Although its accuracy is closing the gap with that of human level under certain settings, one area that can further improve is to incorporate user-specific…

计算与语言 · 计算机科学 2020-05-05 Young Mo Kang , Yingbo Zhou