English
Related papers

Related papers: MotionMERGE: A Multi-granular Framework for Human …

200 papers

Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Jing Lin , Ruisi Wang , Junzhe Lu , Ziqi Huang , Guorui Song , Ailing Zeng , Xian Liu , Chen Wei , Wanqi Yin , Qingping Sun , Zhongang Cai , Lei Yang , Ziwei Liu

We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human-robot group interactions. Effective collaboration in such settings requires consistent situational awareness, built on persistent…

The fusion of Large Language Models with vision models is pioneering new possibilities in user-interactive vision-language tasks. A notable application is reasoning segmentation, where models generate pixel-level segmentation masks by…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Donggon Jang , Yucheol Cho , Suin Lee , Taehyeon Kim , Dae-Shik Kim

As human-robot collaboration advances, natural and flexible communication methods are essential for effective robot control. Traditional methods relying on a single modality or rigid rules struggle with noisy or misaligned data as well as…

Robotics · Computer Science 2025-04-03 Petr Vanc , Karla Stepanova

Despite transformative advances in generative motion synthesis, real-time interactive motion control remains dominated by traditional techniques. In this work, we identify two key challenges in bridging research and production: 1) Real-time…

Motion retrieval is crucial for motion acquisition, offering superior precision, realism, controllability, and editability compared to motion generation. Existing approaches leverage contrastive learning to construct a unified embedding…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Shiyao Yu , Zi-An Wang , Kangning Yin , Zheng Tian , Mingyuan Zhang , Weixin Si , Shihao Zou

Human motion modeling traditionally separates motion generation and estimation into distinct tasks with specialized models. Motion generation models focus on creating diverse, realistic motions from inputs like text, audio, or keyframes,…

Graphics · Computer Science 2025-05-05 Jiefeng Li , Jinkun Cao , Haotian Zhang , Davis Rempe , Jan Kautz , Umar Iqbal , Ye Yuan

Existing 3D human motion generation and understanding methods often exhibit limited interpretability, restricting effective mutual enhancement between these inherently related tasks. While current unified frameworks based on large language…

Artificial Intelligence · Computer Science 2026-01-21 Guocun Wang , Kenkun Liu , Jing Lin , Guorui Song , Jian Li , Xiaoguang Han

We present a unified perspective on tackling various human-centric video tasks by learning human motion representations from large-scale and heterogeneous data resources. Specifically, we propose a pretraining stage in which a motion…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Wentao Zhu , Xiaoxuan Ma , Zhaoyang Liu , Libin Liu , Wayne Wu , Yizhou Wang

In this paper, we tackle the problem of how to build and benchmark a large motion model (LMM). The ultimate goal of LMM is to serve as a foundation model for versatile motion-related tasks, e.g., human motion generation, with…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Liang Xu , Shaoyang Hua , Zili Lin , Yifan Liu , Feipeng Ma , Yichao Yan , Xin Jin , Xiaokang Yang , Wenjun Zeng

Large Language Models(LLMs) have shown remarkable emergent abilities in unifying almost all (if not every) NLP tasks. In the human motion-related realm, however, researchers still develop siloed models for each task. Inspired by InstuctGPT,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Zixiang Zhou , Yu Wan , Baoyuan Wang

Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities. Model merging provides a cost-effective mechanism…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Zibo Shao , Baochen Xiong , Xiaoshan Yang , Yaguang Song , Qimeng Zhang , Haifeng Chen , Changsheng Xu

Despite advancements in Multimodal Large Language Models (MLLMs), their proficiency in fine-grained video motion understanding remains critically limited. They often lack inter-frame differencing and tend to average or ignore subtle visual…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Yipeng Du , Tiehan Fan , Kepan Nan , Rui Xie , Penghao Zhou , Xiang Li , Jian Yang , Zhenheng Yang , Ying Tai

Generative depth estimation methods leverage the rich visual priors stored in pre-trained text-to-image diffusion models, demonstrating astonishing zero-shot capability. However, parameter updates during training lead to catastrophic…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Hongkai Lin , Dingkang Liang , Mingyang Du , Xin Zhou , Xiang Bai

Human motion generation is essential for fields such as animation, robotics, and virtual reality, requiring models that effectively capture motion dynamics from text descriptions. Existing approaches often rely on Contrastive Language-Image…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Gabriel Maldonado , Armin Danesh Pazho , Ghazal Alinezhad Noghre , Vinit Katariya , Hamed Tabkhi

The development of generative artificial intelligence for human motion generation has expanded rapidly, necessitating a unified evaluation framework. This paper presents a detailed review of eight evaluation metrics for human motion…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Ali Ismail-Fawaz , Maxime Devanne , Stefano Berretti , Jonathan Weber , Germain Forestier

Multi-modal Retrieval-Augmented Generation (MMRAG) has emerged as a powerful paradigm for enhancing Multimodal Large Language Models in knowledge-intensive question answering by integrating external visual, textual, and structural…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xiang Fang , Wanlong Fang , Changshuo Wang

3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work shows that 2D videos provide rich, temporally coherent…

Graphics · Computer Science 2026-05-20 Yi-Yang Zhang , Tengjiao Sun , Pengcheng Fang , Deng-Bao Wang , Xiaohao Cai , Min-Ling Zhang , Hansung Kim

We introduce Unimotion, the first unified multi-task human motion model capable of both flexible motion control and frame-level motion understanding. While existing works control avatar motion with global text conditioning, or with…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Chuqiao Li , Julian Chibane , Yannan He , Naama Pearl , Andreas Geiger , Gerard Pons-moll

This study delves into the realm of multi-modality (i.e., video and motion modalities) human behavior understanding by leveraging the powerful capabilities of Large Language Models (LLMs). Diverging from recent LLMs designed for video-only…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Ling-Hao Chen , Shunlin Lu , Ailing Zeng , Hao Zhang , Benyou Wang , Ruimao Zhang , Lei Zhang