English

M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

Computer Vision and Pattern Recognition 2024-11-05 v5

Abstract

This paper presents M3^3GPT, an advanced M\textbf{M}ultimodal, M\textbf{M}ultitask framework for M\textbf{M}otion comprehension and generation. M3^3GPT operates on three fundamental principles. The first focuses on creating a unified representation space for various motion-relevant modalities. We employ discrete vector quantization for multimodal conditional signals, such as text, music and motion/dance, enabling seamless integration into a large language model (LLM) with a single vocabulary. The second involves modeling motion generation directly in the raw motion space. This strategy circumvents the information loss associated with a discrete tokenizer, resulting in more detailed and comprehensive motion generation. Third, M3^3GPT learns to model the connections and synergies among various motion-relevant tasks. Text, the most familiar and well-understood modality for LLMs, is utilized as a bridge to establish connections between different motion tasks, facilitating mutual reinforcement. To our knowledge, M3^3GPT is the first model capable of comprehending and generating motions based on multiple signals. Extensive experiments highlight M3^3GPT's superior performance across various motion-relevant tasks and its powerful zero-shot generalization capabilities for extremely challenging tasks. Project page: \url{https://github.com/luomingshuang/M3GPT}.

Keywords

Cite

@article{arxiv.2405.16273,
  title  = {M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation},
  author = {Mingshuang Luo and Ruibing Hou and Zhuo Li and Hong Chang and Zimo Liu and Yaowei Wang and Shiguang Shan},
  journal= {arXiv preprint arXiv:2405.16273},
  year   = {2024}
}

Comments

Accepted at NeurIPS 2024, 21 pages, 7 figures

R2 v1 2026-06-28T16:40:18.614Z