English
Related papers

Related papers: Unimotion: Unifying 3D Human Motion Synthesis and …

200 papers

Current human motion synthesis frameworks rely on global action descriptions, creating a modality gap that limits both motion understanding and generation capabilities. A single coarse description, such as run, fails to capture details such…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Pengfei Zhang , Pinxin Liu , Pablo Garrido , Hyeongwoo Kim , Bindita Chaudhuri

Motion time series collected from mobile and wearable devices such as smartphones and smartwatches offer significant insights into human behavioral patterns, with wide applications in healthcare, automation, IoT, and AR/XR due to their…

Signal Processing · Electrical Eng. & Systems 2024-10-29 Xiyuan Zhang , Diyan Teng , Ranak Roy Chowdhury , Shuheng Li , Dezhi Hong , Rajesh K. Gupta , Jingbo Shang

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Jingyun Liang , Jingkai Zhou , Shikai Li , Chenjie Cao , Lei Sun , Yichen Qian , Weihua Chen , Fan Wang

Motions in a video primarily consist of camera motion, induced by camera movement, and object motion, resulting from object movement. Accurate control of both camera and object motion is essential for video generation. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Zhouxia Wang , Ziyang Yuan , Xintao Wang , Tianshui Chen , Menghan Xia , Ping Luo , Ying Shan

Recent advances in 3D vision have led to specialized models for either 3D understanding (e.g., shape classification, segmentation, reconstruction) or 3D generation (e.g., synthesis, completion, and editing). However, these tasks are often…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Peng Huang , Yifeng Chen , Zeyu Zhang , Hao Tang

Text-conditioned motion synthesis has made remarkable progress with the emergence of diffusion models. However, the majority of these motion diffusion models are primarily designed for a single character and overlook multi-human…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Zhenzhi Wang , Jingbo Wang , Yixuan Li , Dahua Lin , Bo Dai

The fashion domain encompasses a variety of real-world multimodal tasks, including multimodal retrieval and multimodal generation. The rapid advancements in artificial intelligence generated content, particularly in technologies like large…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Xiangyu Zhao , Yuehan Zhang , Wenlong Zhang , Xiao-Ming Wu

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: the model, the tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Chi Zhang , Jiepeng Wang , Youming Wang , Yuanzhi Liang , Xiaoyan Yang , Zuoxin Li , Haibin Huang , Xuelong Li

Despite the recent progress in text-to-video generation, existing studies usually overlook the issue that only spatial contents but not temporal motions in synthesized videos are under the control of text. Towards such a challenge, this…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Xi Chen , Zhiheng Liu , Mengting Chen , Yutong Feng , Yu Liu , Yujun Shen , Hengshuang Zhao

Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets and paired text descriptions. However, how to effectively…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Xiaoyan Cong , Zekun Li , Zhiyang Dou , Hongyu Li , Omid Taheri , Chuan Guo , Abhay Mittal , Sizhe An , Taku Komura , Wojciech Matusik , Michael J. Black , Srinath Sridhar

We present a novel approach named OmniControl for incorporating flexible spatial control signals into a text-conditioned human motion generation model based on the diffusion process. Unlike previous methods that can only control the pelvis…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Yiming Xie , Varun Jampani , Lei Zhong , Deqing Sun , Huaizu Jiang

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Gaojie Lin , Jianwen Jiang , Jiaqi Yang , Zerong Zheng , Chao Liang

Treating human motion and camera trajectory generation separately overlooks a core principle of cinematography: the tight interplay between actor performance and camera work in the screen space. In this paper, we are the first to cast this…

Graphics · Computer Science 2026-04-02 Robin Courant , Xi Wang , David Loiseaux , Marc Christie , Vicky Kalogeiton

Deep learning has shown significant value in medical image registration for motion correction, however, current techniques are either limited by the type and range of motion they can handle, or require iterative inference and/or retraining…

Image and Video Processing · Electrical Eng. & Systems 2026-05-04 Jian Wang , Razieh Faghihpirayesh , Danny Joca , Polina Golland , Ali Gholipour

Inspired by the strong ties between vision and language, the two intimate human sensing and communication modalities, our paper aims to explore the generation of 3D human full-body motions from texts, as well as its reciprocal task,…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Chuan Guo , Xinxin Zuo , Sen Wang , Li Cheng

Text-driven motion generation offers a powerful and intuitive way to create human movements directly from natural language. By removing the need for predefined motion inputs, it provides a flexible and accessible approach to controlling…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Ali Rida Sahili , Najett Neji , Hedi Tabia

Large language models (LLMs) are, by design, inherently capable of multi-task learning: through a unified next-token prediction paradigm, they can naturally address a wide variety of downstream tasks. Prior work in the motion domain has…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Zeyu Ling , Bo Han , Shiyang Li , Jikang Cheng , Hongdeng Shen , Changqing Zou

Generative modeling of human motion has broad applications in computer animation, virtual reality, and robotics. Conventional approaches develop separate models for different motion synthesis tasks, and typically use a model of a small size…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Jianxin Ma , Shuai Bai , Chang Zhou

We present X-UniMotion, a unified and expressive implicit latent representation for whole-body human motion, encompassing facial expressions, body poses, and hand gestures. Unlike prior motion transfer methods that rely on explicit skeletal…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Guoxian Song , Hongyi Xu , Xiaochen Zhao , You Xie , Tianpei Gu , Zenan Li , Chenxu Zhang , Linjie Luo

A long-standing objective in humanoid robotics is the realization of versatile agents capable of following diverse multimodal instructions with human-level flexibility. Despite advances in humanoid control, bridging high-level multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Nan Jiang , Zimo He , Wanhe Yu , Lexi Pang , Yunhao Li , Hongjie Li , Jieming Cui , Yuhan Li , Yizhou Wang , Yixin Zhu , Siyuan Huang