When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation
Abstract
Raw cosine in the 768-dimensional output space of the Contrastive Style Descriptor (CSD) is now widely read as an absolute, calibrated style-fidelity score for text-to-image and style-imitation evaluation. We introduce the discrimination gap, a corpus-internal, prototype-free and threshold-free diagnostic that tests whether contrastive style cosines admit an absolute same-versus-different interpretation on a candidate artist corpus. On a 1799-artwork, 91-artist public-domain corpus, raw CSD cosine yields negative point-estimate gaps for artists at the pairwise level ( robust under bootstrap) and for in the aggregated-pool scoring regime style-fidelity evaluations typically use. CSLS readout on the frozen backbone reduces the aggregated negative-gap count to ; combined with positional-embedding interpolation to pixels it raises unsupervised pair-verification AUC from to across artist-disjoint splits. We refer to this diagnostic-driven readout protocol on the frozen backbone (CSLS as default, pos-interp as the stronger optional setting) as CSD+, not a new encoder.A cross-backbone check on CLIP-ViT-L/14, SigLIP-large and DINOv2-Large reproduces the same shared-tradition failure pattern, providing evidence that the residual reflects a shared limitation of the four backbones we tested rather than a CSD-specific artefact. Practical implication: before reporting CSD cosine as an absolute style-fidelity score, run the diagnostic on the candidate corpus; CSLS is the minimal correction when it fails.
Cite
@article{arxiv.2605.09030,
title = {When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation},
author = {Jörg Frochte},
journal= {arXiv preprint arXiv:2605.09030},
year = {2026}
}
Comments
24 pages, 7 figures, 19 tables