English
Related papers

Related papers: T3M: Text Guided 3D Human Motion Synthesis from Sp…

200 papers

Given a series of natural language descriptions, our task is to generate 3D human motions that correspond semantically to the text, and follow the temporal order of the instructions. In particular, our goal is to enable the synthesis of a…

Computer Vision and Pattern Recognition · Computer Science 2022-09-13 Nikos Athanasiou , Mathis Petrovich , Michael J. Black , Gül Varol

Audio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce gestures that are…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Xukun Zhou , Fengxin Li , Ming Chen , Yan Zhou , Pengfei Wan , Di Zhang , Yeying Jin , Zhaoxin Fan , Hongyan Liu , Jun He

3D human motion generation has seen substantial advancement in recent years. While state-of-the-art approaches have improved performance significantly, they still struggle with complex and detailed motions unseen in training data, largely…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Shanlin Sun , Gabriel De Araujo , Jiaqi Xu , Shenghan Zhou , Hanwen Zhang , Ziheng Huang , Chenyu You , Xiaohui Xie

Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches assume identity-neutral motion and generate movements using…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Wenqi Jia , Zekun Li , Abhay Mittal , Chengcheng Tang , Chuan Guo , Lezi Wang , James Matthew Rehg , Lingling Tao , Size An

Text-driven human motion generation has recently attracted considerable attention, allowing models to generate human motions based on textual descriptions. However, current methods neglect the influence of human attributes-such as age,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Xinghan Wang , Kun Xu , Fei Li , Cao Sheng , Jiazhong Yu , Yadong Mu

Enabling humanoid robots to synthesize complex, physically coherent motions from natural language commands is a cornerstone of autonomous robotics and human-robot interaction. While diffusion models have shown promise in this text-to-motion…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Wenshuo Chen , Haozhe Jia , Songning Lai , Lei Wang , Yuqi Lin , Hongru Xiao , Lijie Hu , Yutao Yue

Text-to-motion (T2M) generation is becoming a practical tool for animation and interactive avatars. However, modifying specific body parts while maintaining overall motion coherence remains challenging. Existing methods typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Minyue Dai , Ke Fan , Anyi Rao , Jingbo Wang , Bo Dai

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Zhixue Fang , Xu He , Songlin Tang , Haoxian Zhang , Qingfeng Li , Xiaoqiang Liu , Pengfei Wan , Kun Gai

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Youngjoon Jang , Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

Generating varied scenarios through simulation is crucial for training and evaluating safety-critical systems, such as autonomous vehicles. Yet, the task of modeling the trajectories of other vehicles to simulate diverse and meaningful…

Robotics · Computer Science 2024-06-07 Phat Nguyen , Tsun-Hsuan Wang , Zhang-Wei Hong , Sertac Karaman , Daniela Rus

Speech synthesis has significantly advanced from statistical methods to deep neural network architectures, leading to various text-to-speech (TTS) models that closely mimic human speech patterns. However, capturing nuances such as emotion…

Sound · Computer Science 2025-01-14 Shaozuo Zhang , Ambuj Mehrish , Yingting Li , Soujanya Poria

Recent advances in generative modeling and tokenization have driven significant progress in text-to-motion generation, leading to enhanced quality and realism in generated motions. However, effectively leveraging textual information for…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Che-Jui Chang , Qingze Tony Liu , Honglu Zhou , Vladimir Pavlovic , Mubbasir Kapadia

We introduce Unimotion, the first unified multi-task human motion model capable of both flexible motion control and frame-level motion understanding. While existing works control avatar motion with global text conditioning, or with…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Chuqiao Li , Julian Chibane , Yannan He , Naama Pearl , Andreas Geiger , Gerard Pons-moll

Text-to-motion generation, which translates textual descriptions into human motions, has been challenging in accurately capturing detailed user-imagined motions from simple text inputs. This paper introduces StickMotion, an efficient…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Tao Wang , Zhihua Wu , Qiaozhi He , Jiaming Chu , Ling Qian , Yu Cheng , Junliang Xing , Jian Zhao , Lei Jin

Current video generation models usually convert signals indicating appearance and motion received from inputs (e.g., image, text) or latent spaces (e.g., noise vectors) into consecutive frames, fulfilling a stochastic generation process for…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Xue Song , Jingjing Chen , Bin Zhu , Yu-Gang Jiang

We present a framework for video-driven crowd synthesis. Motion vectors extracted from input crowd video are processed to compute global motion paths. These paths encode the dominant motions observed in the input video. These paths are then…

Computer Vision and Pattern Recognition · Computer Science 2018-03-15 Jordan Stadler , Faisal Z. Qureshi

We revisit human motion synthesis, a task useful in various real world applications, in this paper. Whereas a number of methods have been developed previously for this task, they are often limited in two aspects: focusing on the poses while…

Computer Vision and Pattern Recognition · Computer Science 2021-06-01 Jingbo Wang , Sijie Yan , Bo Dai , Dahua LIn

Text-to-speech (TTS) synthesis is a technology that converts written text into spoken words, enabling a natural and accessible means of communication. This abstract explores the key aspects of TTS synthesis, encompassing its underlying…

Software Engineering · Computer Science 2024-01-26 Harini s , Manoj G M

Generating reasonable and high-quality human interactive motions in a given dynamic environment is crucial for understanding, modeling, transferring, and applying human behaviors to both virtual and physical robots. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Peishan Cong , Ziyi Wang , Yuexin Ma , Xiangyu Yue

Recent years have seen an explosion of work and interest in text-to-3D shape generation. Much of the progress is driven by advances in 3D representations, large-scale pretraining and representation learning for text and image data enabling…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Han-Hung Lee , Manolis Savva , Angel X. Chang