English
Related papers

Related papers: T3M: Text Guided 3D Human Motion Synthesis from Sp…

200 papers

In this work, we investigate a simple and must-known conditional generative framework based on Vector Quantised-Variational AutoEncoder (VQ-VAE) and Generative Pre-trained Transformer (GPT) for human motion generation from textural…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Jianrong Zhang , Yangsong Zhang , Xiaodong Cun , Shaoli Huang , Yong Zhang , Hongwei Zhao , Hongtao Lu , Xi Shen

Although significant progress has been made in the field of speech-driven 3D facial animation recently, the speech-driven animation of an indispensable facial component, eye gaze, has been overlooked by recent research. This is primarily…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Yixiang Zhuang , Chunshan Ma , Yao Cheng , Xuan Cheng , Jing Liao , Juncong Lin

In this paper, we introduce LGTM, a novel Local-to-Global pipeline for Text-to-Motion generation. LGTM utilizes a diffusion-based architecture and aims to address the challenge of accurately translating textual descriptions into…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Haowen Sun , Ruikun Zheng , Haibin Huang , Chongyang Ma , Hui Huang , Ruizhen Hu

Speech-driven 3D facial animation with accurate lip synchronization has been widely studied. However, synthesizing realistic motions for the entire face during speech has rarely been explored. In this work, we present a joint audio-text…

Computer Vision and Pattern Recognition · Computer Science 2021-12-08 Yingruo Fan , Zhaojiang Lin , Jun Saito , Wenping Wang , Taku Komura

In this paper, we address the challenge of generating realistic 3D human motions for action classes that were never seen during the training phase. Our approach involves decomposing complex actions into simpler movements, specifically those…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Lorenzo Mandelli , Stefano Berretti

Motion synthesis for diverse object categories holds great potential for 3D content creation but remains underexplored due to two key challenges: (1) the lack of comprehensive motion datasets that include a wide range of high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Wonkwang Lee , Jongwon Jeong , Taehong Moon , Hyeon-Jong Kim , Jaehyeon Kim , Gunhee Kim , Byeong-Uk Lee

With the recent success of deep learning algorithms, many researchers have focused on generative models for human motion animation. However, the research community lacks a platform for training and benchmarking various algorithms, and the…

Graphics · Computer Science 2021-12-14 Yizhou Zhao , Wensi Ai , Liang Qiu , Pan Lu , Feng Shi , Tian Han , Song-Chun Zhu

We propose a real-time system for synthesizing gestures directly from speech. Our data-driven approach is based on Generative Adversarial Neural Networks to model the speech-gesture relationship. We utilize the large amount of speaker video…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Manuel Rebol , Christian Gütl , Krzysztof Pietroszek

Text-to-3D synthesis has recently emerged as a new approach to sampling 3D models by adopting pretrained text-to-image models as guiding visual priors. An intriguing but underexplored problem with existing text-to-3D methods is that 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Uy Dieu Tran , Minh Luu , Phong Ha Nguyen , Khoi Nguyen , Binh-Son Hua

Synthesizing 3D human motion plays an important role in many graphics applications as well as understanding human activity. While many efforts have been made on generating realistic and natural human motion, most approaches neglect the…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Jiashun Wang , Huazhe Xu , Jingwei Xu , Sifei Liu , Xiaolong Wang

Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current methods still face challenges in maintaining lip-sync…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Bin Liu , Zhixiang Xiong , Zhifen He , Bo Li

This paper presents a novel framework for speech-driven gesture production, applicable to virtual agents to enhance human-computer interaction. Specifically, we extend recent deep-learning-based, data-driven methods for speech-driven…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Taras Kucherenko , Dai Hasegawa , Naoshi Kaneko , Gustav Eje Henter , Hedvig Kjellström

Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine-grained detail. Although recent approaches can produce complex and natural…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Heng Li , Xiaotong Lin , Ling-An Zeng , Yulei Kang , Shuai Li , Jian-Fang Hu

Text-to-motion synthesis is a crucial task in computer vision. Existing methods are limited in their universality, as they are tailored for single-person or two-person scenarios and can not be applied to generate motions for more…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Ke Fan , Junshu Tang , Weijian Cao , Ran Yi , Moran Li , Jingyu Gong , Jiangning Zhang , Yabiao Wang , Chengjie Wang , Lizhuang Ma

Recent remarkable advances in large-scale text-to-image diffusion models have inspired a significant breakthrough in text-to-3D generation, pursuing 3D content creation solely from a given text prompt. However, existing text-to-3D…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Yang Chen , Yingwei Pan , Yehao Li , Ting Yao , Tao Mei

Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing audio-video generation pipelines typically decompose the…

Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual concepts within the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-30 Hanxin Zhu , Tianyu He , Anni Tang , Junliang Guo , Zhibo Chen , Jiang Bian

In this study, we introduce T2M-HiFiGPT, a novel conditional generative framework for synthesizing human motion from textual descriptions. This framework is underpinned by a Residual Vector Quantized Variational AutoEncoder (RVQ-VAE) and a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Congyi Wang

In the realm of motion generation, the creation of long-duration, high-quality motion sequences remains a significant challenge. This paper presents our groundbreaking work on "Infinite Motion", a novel approach that leverages long text to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Mengtian Li , Chengshuo Zhai , Shengxiang Yao , Zhifeng Xie , Keyu Chen , Yu-Gang Jiang

Recent techniques for text-to-4D generation synthesize dynamic 3D scenes using supervision from pre-trained text-to-video models. However, existing representations for motion, such as deformation models or time-dependent neural…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Sherwin Bahmani , Xian Liu , Wang Yifan , Ivan Skorokhodov , Victor Rong , Ziwei Liu , Xihui Liu , Jeong Joon Park , Sergey Tulyakov , Gordon Wetzstein , Andrea Tagliasacchi , David B. Lindell
‹ Prev 1 4 5 6 7 8 10 Next ›