English

Renderers are Good Zero-Shot Representation Learners: Exploring Diffusion Latents for Metric Learning

Computer Vision and Pattern Recognition 2023-06-21 v1

Abstract

Can the latent spaces of modern generative neural rendering models serve as representations for 3D-aware discriminative visual understanding tasks? We use retrieval as a proxy for measuring the metric learning properties of the latent spaces of Shap-E, including capturing view-independence and enabling the aggregation of scene representations from the representations of individual image views, and find that Shap-E representations outperform those of the classical EfficientNet baseline representations zero-shot, and is still competitive when both methods are trained using a contrative loss. These findings give preliminary indication that 3D-based rendering and generative models can yield useful representations for discriminative tasks in our innately 3D-native world. Our code is available at \url{https://github.com/michaelwilliamtang/golden-retriever}.

Keywords

Cite

@article{arxiv.2306.10721,
  title  = {Renderers are Good Zero-Shot Representation Learners: Exploring Diffusion Latents for Metric Learning},
  author = {Michael Tang and David Shustin},
  journal= {arXiv preprint arXiv:2306.10721},
  year   = {2023}
}
R2 v1 2026-06-28T11:08:28.403Z