English

ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation

Computer Vision and Pattern Recognition 2025-12-03 v1

Abstract

Existing stylized motion generation models have shown their remarkable ability to understand specific style information from the style motion, and insert it into the content motion. However, capturing intra-style diversity, where a single style should correspond to diverse motion variations, remains a significant challenge. In this paper, we propose a clustering-based framework, ClusterStyle, to address this limitation. Instead of learning an unstructured embedding from each style motion, we leverage a set of prototypes to effectively model diverse style patterns across motions belonging to the same style category. We consider two types of style diversity: global-level diversity among style motions of the same category, and local-level diversity within the temporal dynamics of motion sequences. These components jointly shape two structured style embedding spaces, i.e., global and local, optimized via alignment with non-learnable prototype anchors. Furthermore, we augment the pretrained text-to-motion generation model with the Stylistic Modulation Adapter (SMA) to integrate the style features. Extensive experiments demonstrate that our approach outperforms existing state-of-the-art models in stylized motion generation and motion style transfer.

Keywords

Cite

@article{arxiv.2512.02453,
  title  = {ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation},
  author = {Kerui Chen and Jianrong Zhang and Ming Li and Zhonglong Zheng and Hehe Fan},
  journal= {arXiv preprint arXiv:2512.02453},
  year   = {2025}
}
R2 v1 2026-07-01T08:05:09.348Z