English
Related papers

Related papers: TaCo: Targeted Concept Erasure Prevents Non-Linear…

200 papers

Ensuring that neural models used in real-world applications cannot infer sensitive information, such as demographic attributes like gender or race, from text representations is a critical challenge when fairness is a concern. We address…

Machine Learning · Computer Science 2025-08-19 Antoine Saillenfest , Pirmin Lemberger

Concept Erasure, which aims to prevent pretrained text-to-image models from generating content associated with semantic-harmful concepts (i.e., target concepts), is getting increased attention. State-of-the-art methods formulate this task…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Hongxu Chen , Zhen Wang , Taoran Mei , Lin Li , Bowei Zhu , Runshi Li , Long Chen

Recent work on reducing bias in NLP models usually focuses on protecting or isolating information related to a sensitive attribute (like gender or race). However, when sensitive information is semantically entangled with the task…

Computation and Language · Computer Science 2022-10-25 Zexue He , Yu Wang , Julian McAuley , Bodhisattwa Prasad Majumder

Text-to-video diffusion transformers encode semantic information unevenly across model depth, which constrains effective concept erasure. We identify a representational bottleneck, termed concept-layer topological alignment, under which…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Yiwei Xie , Ping Liu , Zheng Zhang

Natural language processing models tend to learn and encode social biases present in the data. One popular approach for addressing such biases is to eliminate encoded information from the model's representations. However, current methods…

Computation and Language · Computer Science 2023-05-18 Shadi Iskander , Kira Radinsky , Yonatan Belinkov

The representation space of neural models for textual data emerges in an unsupervised manner during training. Understanding how those representations encode human-interpretable concepts is a fundamental problem. One prominent approach for…

Machine Learning · Computer Science 2024-09-17 Shauli Ravfogel , Francisco Vargas , Yoav Goldberg , Ryan Cotterell

Neural network models trained on text data have been found to encode undesirable linguistic or sensitive concepts in their representation. Removing such concepts is non-trivial because of a complex relationship between the concept, text…

Machine Learning · Computer Science 2023-06-21 Abhinav Kumar , Chenhao Tan , Amit Sharma

Modern neural models trained on textual data rely on pre-trained representations that emerge without direct supervision. As these representations are increasingly being used in real-world applications, the inability to \emph{control} their…

Machine Learning · Computer Science 2024-12-18 Shauli Ravfogel , Michael Twiton , Yoav Goldberg , Ryan Cotterell

The ability to control for the kinds of information encoded in neural representation has a variety of use cases, especially in light of the challenge of interpreting these models. We present Iterative Null-space Projection (INLP), a novel…

Computation and Language · Computer Science 2020-04-30 Shauli Ravfogel , Yanai Elazar , Hila Gonen , Michael Twiton , Yoav Goldberg

Recent advances in generative models have demonstrated remarkable capabilities in producing high-quality images, but their reliance on large-scale unlabeled data has raised significant safety and copyright concerns. Efforts to address these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Yang Zhang , Er Jin , Yanfei Dong , Yixuan Wu , Philip Torr , Ashkan Khakzar , Johannes Stegmaier , Kenji Kawaguchi

Concept erasure is extensively utilized in image generation to prevent text-to-image models from generating undesired content. Existing methods can effectively erase narrow concepts that are specific and concrete, such as distinct…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yuze Cai , Jiahao Lu , Hongxiang Shi , Yichao Zhou , Hong Lu

Despite the impressive capabilities of generating images, text-to-image diffusion models are susceptible to producing undesirable outputs such as NSFW content and copyrighted artworks. To address this issue, recent studies have focused on…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Tianyun Yang , Juan Cao , Chang Xu

Concept erasure is the task of erasing information about a concept (e.g., gender or race) from a representation set while retaining the maximum possible utility -- information from original representations. Concept erasure is useful in…

Concept erasure in text-to-image diffusion models seeks to remove undesired concepts while preserving overall generative capability. Localized erasure methods aim to restrict edits to the spatial region occupied by the target concept.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Zhuan Shi , Alireza Dehghanpour Farashah , Rik de Vries , Golnoosh Farnadi

How can we control for latent discrimination in predictive models? How can we provably remove it? Such questions are at the heart of algorithmic fairness and its impacts on society. In this paper, we define a new operational fairness…

Machine Learning · Computer Science 2019-02-25 Soheil Ghili , Ehsan Kazemi , Amin Karbasi

Text-to-image generative models have achieved impressive fidelity and diversity, but can inadvertently produce unsafe or undesirable content due to implicit biases embedded in large-scale training datasets. Existing concept erasure methods,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Jun Li , Lizhi Xiong , Ziqiang Li , Weiwei Jiang , Zhangjie Fu , Yong Li , Guo-Sen Xie

Machine learning models often make predictions based on biased features such as gender, race, and other social attributes, posing significant fairness risks, especially in societal applications, such as hiring, banking, and criminal…

Machine Learning · Computer Science 2024-08-28 Yi Zhang , Dongyuan Lu , Jitao Sang

Studies have been conducted to prevent specific concepts from being generated from pretrained text-to-image generative models, achieving concept erasure in various ways. However, the performance evaluation of these studies is still largely…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Masane Fuchi , Tomohiro Takagi

Concept erasure techniques have recently gained significant attention for their potential to remove unwanted concepts from text-to-image models. While these methods often demonstrate promising results in controlled settings, their…

The burgeoning field of Natural Language Processing (NLP) stands at a critical juncture where the integration of fairness within its frameworks has become an imperative. This PhD thesis addresses the need for equity and transparency in NLP…

Computation and Language · Computer Science 2024-10-17 Fanny Jourdan
‹ Prev 1 2 3 10 Next ›