中文

Motion-X++: 大规模多模态 3D 整体人体动作数据集

计算机视觉与模式识别 2025-01-10 v1

摘要

本文介绍了 Motion-X++,一个大规模多模态 3D 表达性整体人体动作数据集。现有动作数据集主要捕获身体姿态,缺乏面部表情、手势以及细粒度姿态描述,且通常限制在实验室环境中,仅能手动标注文本描述,从而限制了其可扩展性。为此,我们开发了可扩展的标注管道,能够自动从 RGB 视频中捕获 3D 整体人体动作及综合纹理标签,构建由 81.1K 条文本-动作配对组成的 Motion-X 数据集。进一步地,我们将 Motion-X 扩展为 Motion-X++,通过改进标注管道、引入更多数据模态并扩大数据规模。Motion-X++ 提供了 19.5 万 3D 整体人体姿态标注,覆盖 120.5K 条动作序列的大规模场景,包含 80.8K 条 RGB 视频、45.3K 条音频、19.5 万帧级整体人体姿态描述以及 120.5K 条序列级语义标签。全面实验验证了标注管道的准确性,并突出 Motion-X++ 在生成带有配对多模态标签的表达、精确且自然的动作方面的显著优势,支持多项下游任务,包括文本驱动的整体人体动作生成、音频驱动的动作生成、3D 整体人体网格恢复以及 2D 整体人体关键点估计等。

关键词

引用

@article{arxiv.2501.05098,
  title  = {Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset},
  author = {Yuhong Zhang and Jing Lin and Ailing Zeng and Guanlin Wu and Shunlin Lu and Yurong Fu and Yuanhao Cai and Ruimao Zhang and Haoqian Wang and Lei Zhang},
  journal= {arXiv preprint arXiv:2501.05098},
  year   = {2025}
}

备注

17 pages, 14 figures, This work extends and enhances the research published in the NeurIPS 2023 paper, "Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset". arXiv admin note: substantial text overlap with arXiv:2307.00818