Spatially Localized Image Degradation Embeddings for Image Quality Assessment
Abstract
Self-supervised learning (SSL) currently drives state-of-the-art performance in no-reference image quality assessment (NR-IQA). However, standard SSL pipelines uniformly apply synthetic distortions across the entire image field, which can limit their sensitivity to spatially localized and co-occurring degradations encountered in real-world content. In this work, we empirically expose this representational blind spot across existing state-of-the-art encoders, demonstrating their reduced sensitivity to spatially bounded image degradations. To bridge this gap, we introduce Spatial Localized Image Degradation Embeddings for Image Quality Assessment (SLIDE-IQA). SLIDE-IQA employs a dual-branch Vision Transformer framework that injects spatially bounded degradations into a contrastive pretraining objective. To handle the spatial complexity of these degradations, we introduce a Threshold-Bounded Exclusion Mechanism, a representational design choice that resolves structural conflicts arising from spatially localized distortions to ensure the latent space respects both degradation type and spatial scale. Finally, we show that SLIDE-IQA's synthetic-only pretraining significantly improves sensitivity to localized distortions, while achieving competitive performance on NR-IQA benchmarks against existing SSL NR-IQA models.
Keywords
Cite
@article{arxiv.2606.29162,
title = {Spatially Localized Image Degradation Embeddings for Image Quality Assessment},
author = {Krishna Srikar Durbha and Hassene Tmar and Ping-Hao Wu and Ioannis Katsavounidis and Alan C. Bovik},
journal= {arXiv preprint arXiv:2606.29162},
year = {2026}
}
Comments
Under Review