English

Twist and Compute: The Cost of Pose in 3D Generative Diffusion

Computer Vision and Pattern Recognition 2025-11-12 v1

Abstract

Despite their impressive results, large-scale image-to-3D generative models remain opaque in their inductive biases. We identify a significant limitation in image-conditioned 3D generative models: a strong canonical view bias. Through controlled experiments using simple 2D rotations, we show that the state-of-the-art Hunyuan3D 2.0 model can struggle to generalize across viewpoints, with performance degrading under rotated inputs. We show that this failure can be mitigated by a lightweight CNN that detects and corrects input orientation, restoring model performance without modifying the generative backbone. Our findings raise an important open question: Is scale enough, or should we pursue modular, symmetry-aware designs?

Keywords

Cite

@article{arxiv.2511.08203,
  title  = {Twist and Compute: The Cost of Pose in 3D Generative Diffusion},
  author = {Kyle Fogarty and Jack Foster and Boqiao Zhang and Jing Yang and Cengiz Öztireli},
  journal= {arXiv preprint arXiv:2511.08203},
  year   = {2025}
}

Comments

Accepted to EurIPS 2025 Workshop on Principles of Generative Modeling (PriGM)