HY-Motion 1.0:规模化的流匹配模型用于文本到运动生成
摘要
我们提出 HY-Motion 1.0, 一系列具有最先进水平、大规模的运动生成模型, capable of generating 3D human motions from textual descriptions. HY-Motion 1.0 represents the first successful attempt to scale up Diffusion Transformer (DiT)-based flow matching models to the billion-parameter scale within the motion generation domain, delivering instruction-following capabilities that significantly outperform current open-source benchmarks. Uniquely, we introduce a comprehensive, full-stage training paradigm -- including large-scale pretraining on over 3,000 hours of motion data, high-quality fine-tuning on 400 hours of curated data, and reinforcement learning from both human feedback and reward models -- to ensure precise alignment with the text instruction and high motion quality. This framework is supported by our meticulous data processing pipeline, which performs rigorous motion cleaning and captioning. Consequently, our model achieves the most extensive coverage, spanning over 200 motion categories across 6 major classes. We release HY-Motion 1.0 to the open-source community to foster future research and accelerate the transition of 3D human motion generation models towards commercial maturity.
关键词
引用
@article{arxiv.2512.23464,
title = {HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation},
author = {Yuxin Wen and Qing Shuai and Di Kang and Jing Li and Cheng Wen and Yue Qian and Ningxin Jiao and Changhai Chen and Weijie Chen and Yiran Wang and Jinkun Guo and Dongyue An and Han Liu and Yanyu Tong and Chao Zhang and Qing Guo and Juan Chen and Qiao Zhang and Youyi Zhang and Zihao Yao and Cheng Zhang and Hong Duan and Xiaoping Wu and Qi Chen and Fei Cheng and Liang Dong and Peng He and Hao Zhang and Jiaxin Lin and Chao Zhang and Zhongyi Fan and Yifan Li and Zhichao Hu and Yuhong Liu and Linus and Jie Jiang and Xiaolong Li and Linchao Bao},
journal= {arXiv preprint arXiv:2512.23464},
year = {2025}
}
备注
Github: see https://github.com/Tencent-Hunyuan/HY-Motion-1.0