English
Related papers

Related papers: T3M: Text Guided 3D Human Motion Synthesis from Sp…

200 papers

Recent works have sought to enhance the controllability and precision of text-driven motion generation. Some approaches leverage large language models (LLMs) to produce more detailed texts, while others incorporate global 3D coordinate…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Keming Shen , Bizhu Wu , Junliang Chen , Xiaoqin Wang , Linlin Shen

Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and robotics. However, this task faces critical challenges: the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Wanjiang Weng , Xiaofeng Tan , Hongsong Wang , Pan Zhou

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

Speech-driven 3D facial animation aims to synthesize realistic facial motion sequences from given audio, matching the speaker's speaking style. However, previous works often require priors such as class labels of a speaker or additional 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Hyung Kyu Kim , Sangmin Lee , Hak Gu Kim

Denoising diffusion models have shown great promise in human motion synthesis conditioned on natural language descriptions. However, integrating spatial constraints, such as pre-defined motion trajectories and obstacles, remains a challenge…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Korrawe Karunratanakul , Konpat Preechakul , Supasorn Suwajanakorn , Siyu Tang

Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characteristics of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality…

Computer Vision and Pattern Recognition · Computer Science 2022-10-19 Zan Wang , Yixin Chen , Tengyu Liu , Yixin Zhu , Wei Liang , Siyuan Huang

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized Variational…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Sohan Anisetty , James Hays

Human video synthesis aims to create lifelike characters in various environments, with wide applications in VR, storytelling, and content creation. While 2D diffusion-based methods have made significant progress, they struggle to generalize…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Liyuan Cui , Xiaogang Xu , Wenqi Dong , Zesong Yang , Hujun Bao , Zhaopeng Cui

Controllable generation of 3D human motions becomes an important topic as the world embraces digital transformation. Existing works, though making promising progress with the advent of diffusion models, heavily rely on meticulously captured…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Nhat M. Hoang , Kehong Gong , Chuan Guo , Michael Bi Mi

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts describing the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Yuming Jiang , Shuai Yang , Tong Liang Koh , Wayne Wu , Chen Change Loy , Ziwei Liu

3D functionality segmentation aims to identify the interactive element in a 3D scene required to perform an action described in free-form language (e.g., the handle to ``Open the second drawer of the cabinet near the bed''). Progress has…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Jaime Corsetti , Francesco Giuliari , Davide Boscaini , Pedro Hermosilla , Andrea Pilzer , Guofeng Mei , Alexandros Delitzas , Francis Engelmann , Fabio Poiesi

3D content creation plays a vital role in various applications, such as gaming, robotics simulation, and virtual reality. However, the process is labor-intensive and time-consuming, requiring skilled designers to invest considerable effort…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Chenhan Jiang

Text-to-motion generation has experienced remarkable progress in recent years. However, current approaches remain limited to synthesizing motion from short or general text prompts, primarily due to dataset constraints. This limitation…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Chuan Guo , Inwoo Hwang , Jian Wang , Bing Zhou

In this paper, we introduce DirectorLLM, a novel video generation model that employs a large language model (LLM) to orchestrate human poses within videos. As foundational text-to-video models rapidly evolve, the demand for high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Kunpeng Song , Tingbo Hou , Zecheng He , Haoyu Ma , Jialiang Wang , Animesh Sinha , Sam Tsai , Yaqiao Luo , Xiaoliang Dai , Li Chen , Xide Xia , Peizhao Zhang , Peter Vajda , Ahmed Elgammal , Felix Juefei-Xu

We present a novel approach for synthesizing 3D facial motions from audio sequences using key motion embeddings. Despite recent advancements in data-driven techniques, accurately mapping between audio signals and 3D facial meshes remains…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Zhihao Xu , Shengjie Gong , Jiapeng Tang , Lingyu Liang , Yining Huang , Haojie Li , Shuangping Huang

Recent advancements in models linking natural language with human motions have shown significant promise in motion generation and editing based on instructional text. Motivated by applications in sports coaching and motor skill learning, we…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Qihang Fang , Chengcheng Tang , Bugra Tekin , Yanchao Yang

Generating realistic 3D hand motion from natural language is vital for VR, robotics, and human-computer interaction. Existing methods either focus on full-body motion, overlooking detailed hand gestures, or require explicit 3D object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Ching-Lam Cheng , Bin Zhu , Shengfeng He

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output animations from speech…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Kiran Chhatre , Radek Daněček , Nikos Athanasiou , Giorgio Becherini , Christopher Peters , Michael J. Black , Timo Bolkart

Although humans engaged in face-to-face conversation simultaneously communicate both verbally and non-verbally, methods for joint and unified synthesis of speech audio and co-speech 3D gesture motion from text are a new and emerging field.…

Human-Computer Interaction · Computer Science 2024-05-01 Shivam Mehta , Anna Deichler , Jim O'Regan , Birger Moëll , Jonas Beskow , Gustav Eje Henter , Simon Alexanderson

People communicate using both speech and non-verbal signals such as gestures, face expression or body pose. Non-verbal signals impact the meaning of the spoken utterance in an abundance of ways. An absence of non-verbal signals impoverishes…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Manuel Rebol , Christian Gütl , Krzysztof Pietroszek