English
Related papers

Related papers: Safety and Fairness for Content Moderation in Gene…

200 papers

Threat modelling is the process of identifying potential vulnerabilities in a system and prioritising them. Existing threat modelling tools focus primarily on technical systems and are not as well suited to interpersonal threats. In this…

Cryptography and Security · Computer Science 2025-05-06 Kieron Ivy Turk , Anna Talas , Alice Hutchings

Ensuring equitable Artificial Intelligence (AI) in healthcare demands systems that make unbiased decisions across all demographic groups, bridging technical innovation with ethical principles. Foundation Models (FMs), trained on vast…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Dilermando Queiroz , Anderson Carlos , André Anjos , Lilian Berton

Large generative AI models (GMs) like GPT and DALL-E are trained to generate content for general, wide-ranging purposes. GM content filters are generalized to filter out content which has a risk of harm in many cases, e.g., hate speech.…

Human-Computer Interaction · Computer Science 2023-06-07 Logan Stapleton , Jordan Taylor , Sarah Fox , Tongshuang Wu , Haiyi Zhu

Recent progress in generative AI, especially diffusion models, has demonstrated significant utility in text-to-image synthesis. Particularly in healthcare, these models offer immense potential in generating synthetic datasets and training…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Yan Luo , Muhammad Osama Khan , Congcong Wen , Muhammad Muneeb Afzal , Titus Fidelis Wuermeling , Min Shi , Yu Tian , Yi Fang , Mengyu Wang

The exponential growth of social media platforms has brought about a revolution in communication and content dissemination in human society. Nevertheless, these platforms are being increasingly misused to spread toxic content, including…

Software Engineering · Computer Science 2023-08-22 Wenxuan Wang , Jingyuan Huang , Jen-tse Huang , Chang Chen , Jiazhen Gu , Pinjia He , Michael R. Lyu

While data-driven predictive models are a strictly technological construct, they may operate within a social context in which benign engineering choices entail implicit, indirect and unexpected real-life consequences. Fairness of such…

Machine Learning · Computer Science 2024-07-11 Kacper Sokol , Meelis Kull , Jeffrey Chan , Flora Salim

Recent work has shown that a model's input word embeddings can serve as effective control variables for steering its behavior toward outputs that satisfy desired properties. However, this has only been demonstrated for pretrained…

Computation and Language · Computer Science 2026-04-30 Baturay Saglam , Dionysis Kalogerias

We present a framework for the automated measurement of responsible AI (RAI) metrics for large language models (LLMs) and associated products and services. Our framework for automatically measuring harms from LLMs builds on existing…

Many generative foundation models (or GFMs) are trained on publicly available data and use public infrastructure, but 1) may degrade the "digital commons" that they depend on, and 2) do not have processes in place to return value captured…

Computers and Society · Computer Science 2023-03-21 Saffron Huang , Divya Siddarth

Motivated by ethical and legal concerns, the scientific community is actively developing methods to limit the misuse of Text-to-Image diffusion models for reproducing copyrighted, violent, explicit, or personal information in the generated…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Vitali Petsiuk , Kate Saenko

The ingrained principles of fairness in a dialogue system's decision-making process and generated responses are crucial for user engagement, satisfaction, and task achievement. Absence of equitable and inclusive principles can hinder the…

Computation and Language · Computer Science 2023-07-11 Anthony Sicilia , Malihe Alikhani

Previous work has largely considered the fairness of image captioning systems through the underspecified lens of "bias." In contrast, we present a set of techniques for measuring five types of representational harms, as well as the…

Computers and Society · Computer Science 2022-06-16 Angelina Wang , Solon Barocas , Kristen Laird , Hanna Wallach

Generative AI systems are transforming content creation, but their usability remains a key challenge. This paper examines usability factors such as user experience, transparency, control, and cognitive load. Common challenges include…

Human-Computer Interaction · Computer Science 2025-02-26 Anna Ravera , Cristina Gena

Generative artificial intelligence (AI) is a widely popular technology that will have a profound impact on society and individuals. Less than a decade ago, it was thought that creative work would be among the last to be automated - yet…

Human-Computer Interaction · Computer Science 2023-08-21 Jonas Oppenlaender , Johanna Silvennoinen , Ville Paananen , Aku Visuri

Accountability aims to provide explanations for why unwanted situations occurred, thus providing means to assign responsibility and liability. As such, accountability has slightly different meanings across the sciences. In computer science,…

Computers and Society · Computer Science 2016-08-30 Severin Kacianka , Florian Kelbert , Alexander Pretschner

It has been shown that many generative models inherit and amplify societal biases. To date, there is no uniform/systematic agreed standard to control/adjust for these biases. This study examines the presence and manipulation of societal…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Philip Wootaek Shin , Jihyun Janice Ahn , Wenpeng Yin , Jack Sampson , Vijaykrishnan Narayanan

As generative AI systems become widely adopted, they enable unprecedented creation levels of synthetic data across text, images, audio, and video modalities. While research has addressed the energy consumption of model training and…

Computers and Society · Computer Science 2025-05-29 Vanessa Utz

Social platforms have revolutionized information sharing, but also accelerated the dissemination of harmful and policy-violating content. To ensure safety and compliance at scale, moderation systems must go beyond efficiency and offer…

Computation and Language · Computer Science 2026-01-09 Anqi Li , Wenwei Jin , Jintao Tong , Pengda Qin , Weijia Li , Guo Lu

Although automated harmful content detection systems are frequently used to monitor online platforms, moderators and end users frequently cannot understand the logic underlying their predictions. While recent studies have focused on…

Computation and Language · Computer Science 2026-03-20 Trishita Dhara , Siddhesh Sheth

The proliferation of generative models, such as Generative Adversarial Networks (GANs), Diffusion Models, and Variational Autoencoders (VAEs), has enabled the synthesis of high-quality multimedia data. However, these advancements have also…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Arpan Mahara , Naphtali Rishe