English

$\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models

Computer Vision and Pattern Recognition 2025-08-05 v3

Abstract

We study how\textit{how} rich visual semantic information is represented within various layers and denoising timesteps of different diffusion architectures. We uncover monosemantic interpretable features by leveraging k-sparse autoencoders (k-SAE). We substantiate our mechanistic interpretations via transfer learning using light-weight classifiers on off-the-shelf diffusion models' features. On 44 datasets, we demonstrate the effectiveness of diffusion features for representation learning. We provide an in-depth analysis of how different diffusion architectures, pre-training datasets, and language model conditioning impacts visual representation granularity, inductive biases, and transfer learning capabilities. Our work is a critical step towards deepening interpretability of black-box diffusion models. Code and visualizations available at: https://github.com/revelio-diffusion/revelio

Keywords

Cite

@article{arxiv.2411.16725,
  title  = {$\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models},
  author = {Dahye Kim and Xavier Thomas and Deepti Ghadiyaram},
  journal= {arXiv preprint arXiv:2411.16725},
  year   = {2025}
}

Comments

ICCV 2025

R2 v1 2026-06-28T20:11:59.675Z