中文
相关论文

相关论文: LEACE: Perfect linear concept erasure in closed fo…

200 篇论文

We address the problem of concept removal in deep neural networks, aiming to learn representations that do not encode certain specified concepts (e.g., gender etc.) We propose a novel method based on adversarial linear classifiers trained…

机器学习 · 计算机科学 2023-10-10 Yegor Klochkov , Jean-Francois Ton , Ruocheng Guo , Yang Liu , Hang Li

LLMs have revolutionized the field of artificial intelligence and have emerged as the de-facto tool for many tasks. The current established technology of LLMs is to process input and generate output at the token level. This is in sharp…

The powerful generative capabilities of diffusion models have raised growing privacy and safety concerns regarding generating sensitive or undesired content. In response, machine unlearning (MU) -- commonly referred to as concept erasure…

机器学习 · 计算机科学 2026-03-03 Xinwen Cheng , Jingyuan Zhang , Zhehao Huang , Yingwen Wu , Xiaolin Huang

Accurately aligning contextual representations in cross-lingual sentence embeddings is key for effective parallel data mining. A common strategy for achieving this alignment involves disentangling semantics and language in sentence…

计算与语言 · 计算机科学 2025-09-03 Dayeon Ki , Cheonbok Park , Hyunjoong Kim

Recent advances in text-to-image diffusion models enable photorealistic image generation, but they also risk producing malicious content, such as NSFW images. To mitigate risk, concept erasure methods are studied to facilitate the model to…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Ruidong Chen , Honglin Guo , Lanjun Wang , Chenyu Zhang , Weizhi Nie , An-An Liu

Bias in word embeddings such as Word2Vec has been widely investigated, and many efforts made to remove such bias. We show how to use conceptors debiasing to post-process both traditional and contextualized word embeddings. Our conceptor…

计算与语言 · 计算机科学 2019-06-17 Saket Karve , Lyle Ungar , João Sedoc

Current systems of fine-grained entity typing use distant supervision in conjunction with existing knowledge bases to assign categories (type labels) to entity mentions. However, the type labels so obtained from knowledge bases are often…

计算与语言 · 计算机科学 2016-02-18 Xiang Ren , Wenqi He , Meng Qu , Clare R. Voss , Heng Ji , Jiawei Han

Large-scale diffusion models, known for their impressive image generation capabilities, have raised concerns among researchers regarding social impacts, such as the imitation of copyrighted artistic styles. In response, existing approaches…

机器学习 · 计算机科学 2024-02-12 Mengnan Zhao , Lihe Zhang , Tianhang Zheng , Yuqiu Kong , Baocai Yin

Language representations are known to carry stereotypical biases and, as a result, lead to biased predictions in downstream tasks. While existing methods are effective at mitigating biases by linear projection, such methods are too…

计算与语言 · 计算机科学 2021-09-14 Sunipa Dev , Tao Li , Jeff M Phillips , Vivek Srikumar

The rapid proliferation of large-scale text-to-image diffusion (T2ID) models has raised serious concerns about their potential misuse in generating harmful content. Although numerous methods have been proposed for erasing undesired concepts…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Koushik Srivatsan , Fahad Shamshad , Muzammal Naseer , Vishal M. Patel , Karthik Nandakumar

Text-to-image (T2I) diffusion models often inadvertently generate unwanted concepts such as watermarks and unsafe images. These concepts, termed as the "implicit concepts", could be unintentionally learned during training and then be…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Zhili Liu , Kai Chen , Yifan Zhang , Jianhua Han , Lanqing Hong , Hang Xu , Zhenguo Li , Dit-Yan Yeung , James Kwok

Interpreting individual neurons in deep neural networks is a crucial step towards understanding their complex decision-making processes and ensuring AI safety. Despite recent progress in neuron labeling, existing methods often limit the…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Vladimir Zaigrajew , Michał Piechota , Gaspar Sekula , Paweł Gelar , Przemysław Biecek

Despite the impressive capabilities of generating images, text-to-image diffusion models are susceptible to producing undesirable outputs such as NSFW content and copyrighted artworks. To address this issue, recent studies have focused on…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Tianyun Yang , Juan Cao , Chang Xu

Diffusion models (DMs) have achieved significant progress in text-to-image generation. However, the inevitable inclusion of sensitive information during pre-training poses safety risks, such as unsafe content generation and copyright…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Hongguang Zhu , Yunchao Wei , Mengyu Wang , Siyu Jiao , Yan Fang , Jiannan Huang , Yao Zhao

Diffusion models have transformed image generation, yet controlling their outputs to reliably erase undesired concepts remains challenging. Existing approaches usually require task-specific training and struggle to generalize across both…

Concept erasure techniques for text-to-video (T2V) diffusion models report substantial suppression of sensitive content, yet current evaluation is limited to checking whether the target concept is absent from generated frames, treating…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Yiwei Xie , Zheng Zhang , Ping Liu

The linear subspace hypothesis (Bolukbasi et al., 2016) states that, in a language model's representation space, all information about a concept such as verbal number is encoded in a linear subspace. Prior work has relied on auxiliary…

计算与语言 · 计算机科学 2025-03-27 Clément Guerner , Tianyu Liu , Anej Svete , Alexander Warstadt , Ryan Cotterell

Ensuring fairness in image classification prevents models from perpetuating and amplifying bias. Concept bottleneck models (CBMs) map images to high-level, human-interpretable concepts before making predictions via a sparse, one-layer…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Schrasing Tong , Antoine Salaun , Vincent Yuan , Annabel Adeyeri , Lalana Kagal

Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment alignment and safety settings. However, despite its empirical success, it currently lacks a…

Existing unlearning algorithms in text-to-image generative models often fail to preserve the knowledge of semantically related concepts when removing specific target concepts: a challenge known as adjacency. To address this, we propose FADE…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Kartik Thakral , Tamar Glaser , Tal Hassner , Mayank Vatsa , Richa Singh