We present a full reference, perceptual image metric based on VGG-16, an artificial neural network trained on object classification. We fit the metric to a new database based on 140k unique images annotated with ground truth by human raters who received minimal instruction. The resulting metric shows competitive performance on TID 2013, a database widely used to assess image quality assessments methods. More interestingly, it shows strong responses to objects potentially carrying semantic relevance such as faces and text, which we demonstrate using a visualization technique and ablation experiments. In effect, the metric appears to model a higher influence of semantic context on judgments, which we observe particularly in untrained raters. As the vast majority of users of image processing systems are unfamiliar with Image Quality Assessment (IQA) tasks, these findings may have significant impact on real-world applications of perceptual metrics.
@article{arxiv.1808.00447,
title = {Towards a Semantic Perceptual Image Metric},
author = {Troy Chinen and Johannes Ballé and Chunhui Gu and Sung Jin Hwang and Sergey Ioffe and Nick Johnston and Thomas Leung and David Minnen and Sean O'Malley and Charles Rosenberg and George Toderici},
journal= {arXiv preprint arXiv:1808.00447},
year = {2018}
}