Sapiens:面向人类视觉模型的基础
摘要
我们提出 Sapiens,一个面向四项基本人类视觉任务的模型系列——2D 姿态估计、身体部位分割、深度估计和表面法线预测。我们的模型原生支持 1K 高分辨率推理,并且极其容易通过对 300 万以上野外人类图像进行预训练来适应单个任务。我们观察到,在相同的计算预算下,针对人类图像精选数据集进行自监督预训练显著提升了多样化人类中心任务的性能。 resulting models exhibit remarkable generalization to in-the-wild data, even when labeled data is scarce or entirely synthetic. Our simple model design also brings scalability -- model performance across tasks improves as we scale the number of parameters from 0.3 to 2 billion. Sapiens consistently surpasses existing baselines across various human-centric benchmarks. We achieve significant improvements over the prior state-of-the-art on Humans-5K (pose) by 7.6 mAP, Humans-2K (part-seg) by 17.1 mIoU, Hi4D (depth) by 22.4% relative RMSE, and THuman2 (normal) by 53.5% relative angular error. Project page: https://about.meta.com/realitylabs/codecavatars/sapiens.
引用
@article{arxiv.2408.12569,
title = {Sapiens: Foundation for Human Vision Models},
author = {Rawal Khirodkar and Timur Bagautdinov and Julieta Martinez and Su Zhaoen and Austin James and Peter Selednik and Stuart Anderson and Shunsuke Saito},
journal= {arXiv preprint arXiv:2408.12569},
year = {2024}
}
备注
ECCV 2024 (Oral)