English
Related papers

Related papers: Concept activation vectors: a unifying view and ad…

200 papers

Robustness of machine learning models on ever-changing real-world data is critical, especially for applications affecting human well-being such as content moderation. New kinds of abusive language continually emerge in online discussions in…

Computation and Language · Computer Science 2022-04-06 Isar Nejadgholi , Kathleen C. Fraser , Svetlana Kiritchenko

Humans use abstract concepts for understanding instead of hard features. Recent interpretability research has focused on human-centered concept explanations of neural networks. Concept Activation Vectors (CAVs) estimate a model's…

Machine Learning · Computer Science 2023-11-28 Avani Gupta , Saurabh Saini , P J Narayanan

Machine learning models are trained with relatively simple objectives, such as next token prediction. However, on deployment, they appear to capture a more fundamental representation of their input data. It is of interest to understand the…

Machine Learning · Computer Science 2024-12-23 Thomas Walker

We propose a concept-based adversarial attack framework that extends beyond single-image perturbations by adopting a probabilistic perspective. Rather than modifying a single image, our method operates on an entire concept - represented by…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Andi Zhang , Xuan Ding , Steven McDonagh , Samuel Kaski

Multimodal Emotion Recognition refers to the classification of input video sequences into emotion labels based on multiple input modalities (usually video, audio and text). In recent years, Deep Neural networks have shown remarkable…

Machine Learning · Computer Science 2024-10-28 Ashish Ramayee Asokan , Nidarshan Kumar , Anirudh Venkata Ragam , Shylaja S Sharath

As applications of generative AI become mainstream, it is important to understand what generative models are capable of producing, and the extent to which one can predictably control their outputs. In this paper, we propose a visualization…

Human-Computer Interaction · Computer Science 2024-07-01 Sangwon Jeong , Mingwei Li , Matthew Berger , Shusen Liu

Learning useful representations of complex data has been the subject of extensive research for many years. With the diffusion of Deep Neural Networks, Variational Autoencoders have gained lots of attention since they provide an explicit…

Machine Learning · Computer Science 2020-09-15 Marco Maggipinto , Matteo Terzi , Gian Antonio Susto

We propose to generate adversarial samples by modifying activations of upper layers encoding semantically meaningful concepts. The original sample is shifted towards a target sample, yielding an adversarial sample, by using the modified…

Machine Learning · Computer Science 2022-03-22 Johannes Schneider , Giovanni Apruzzese

Distributed word vector spaces are considered hard to interpret which hinders the understanding of natural language processing (NLP) models. In this work, we introduce a new method to interpret arbitrary samples from a word vector space. To…

Computation and Language · Computer Science 2019-04-03 Robert Schwarzenberg , Lisa Raithel , David Harbecke

Recent research in deep learning methodology has led to a variety of complex modelling techniques in computer vision (CV) that reach or even outperform human performance. Although these black-box deep learning models have obtained…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Anh Pham Thi Minh

Concept vectors aim to enhance model interpretability by linking internal representations with human-understandable semantics, but their utility is often limited by noisy and inconsistent activations. In this work, we uncover a clear…

Machine Learning · Computer Science 2025-12-05 Cassandra Goldberg , Chaehyeon Kim , Adam Stein , Eric Wong

Concept Activation Vectors (CAVs) are widely used to model human-understandable concepts as directions within the latent space of neural networks. They are trained by identifying directions from the activations of concept samples to those…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Eren Erogullari , Sebastian Lapuschkin , Wojciech Samek , Frederik Pahde

Explanations for deep neural network predictions in terms of domain-related concepts can be valuable in medical applications, where justifications are important for confidence in the decision-making. In this work, we propose a methodology…

Machine Learning · Computer Science 2019-04-10 Mara Graziani , Vincent Andrearczyk , Henning Müller

Collective variables (CVs) are low-dimensional projections of high-dimensional system states. They are used to gain insights into complex emergent dynamical behaviors of processes on networks. The relation between CVs and network measures…

Physics and Society · Physics 2026-03-19 Marvin Lücke , Stefanie Winkelmann , Jobst Heitzig , Nora Molkenthin , Péter Koltai

Despite their outstanding accuracy, semi-supervised segmentation methods based on deep neural networks can still yield predictions that are considered anatomically impossible by clinicians, for instance, containing holes or disconnected…

Computer Vision and Pattern Recognition · Computer Science 2021-07-14 Ping Wang , Jizong Peng , Marco Pedersoli , Yuanfeng Zhou , Caiming Zhang , Christian Desrosiers

Interpreting the inner workings of deep learning models is crucial for establishing trust and ensuring model safety. Concept-based explanations have emerged as a superior approach that is more interpretable than feature attribution…

Machine Learning · Computer Science 2023-07-17 Mara Graziani , Laura O' Mahony , An-Phi Nguyen , Henning Müller , Vincent Andrearczyk

Attribution methods for Vision Transformers (ViTs) aim to identify image regions that influence model predictions, but producing faithful and well-localized attributions remains challenging. Existing attribution methods face several…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Amirmohammad Izadi , Mohammadali Banayeeanzade , Alireza Mirrokni , Hosein Hasani , Mobin Bagherian , Faridoun Mehri , Mahdieh Soleymani Baghshah

Collaborative perception, which greatly enhances the sensing capability of connected and autonomous vehicles (CAVs) by incorporating data from external resources, also brings forth potential security risks. CAVs' driving decisions rely on…

Cryptography and Security · Computer Science 2023-10-04 Qingzhao Zhang , Shuowei Jin , Ruiyang Zhu , Jiachen Sun , Xumiao Zhang , Qi Alfred Chen , Z. Morley Mao

Although saliency maps can highlight important regions to explain the reasoning behind image classification in artificial intelligence (AI), the meaning of these regions is left to the user's interpretation. In contrast, conceptbased…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Michihiro Kuroki , Toshihiko Yamasaki

Variational Autoencoders (VAEs) are expressive latent variable models that can be used to learn complex probability distributions from training data. However, the quality of the resulting model crucially relies on the expressiveness of the…

Machine Learning · Computer Science 2018-06-12 Lars Mescheder , Sebastian Nowozin , Andreas Geiger