English
Related papers

Related papers: Text2Performer: Text-Driven Human Video Generation

200 papers

Text animation serves as an expressive medium, transforming static communication into dynamic experiences by infusing words with motion to evoke emotions, emphasize meanings, and construct compelling narratives. Crafting animations that are…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Zichen Liu , Yihao Meng , Hao Ouyang , Yue Yu , Bolin Zhao , Daniel Cohen-Or , Huamin Qu

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions…

Machine Learning · Computer Science 2021-01-15 Simon Alexanderson , Éva Székely , Gustav Eje Henter , Taras Kucherenko , Jonas Beskow

Recent breakthroughs in video generation, powered by large-scale datasets and diffusion techniques, have shown that video diffusion models can function as implicit 4D novel view synthesizers. Nevertheless, current methods primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Yihao Zhi , Chenghong Li , Hongjie Liao , Xihe Yang , Zhengwentai Sun , Jiahao Chang , Xiaodong Cun , Wensen Feng , Xiaoguang Han

Text-to-video generation aims to produce a video based on a given prompt. Recently, several commercial video models have been able to generate plausible videos with minimal noise, excellent details, and high aesthetic scores. However, these…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Haoxin Chen , Yong Zhang , Xiaodong Cun , Menghan Xia , Xintao Wang , Chao Weng , Ying Shan

The generation of humanoid animation from text prompts can profoundly impact animation production and AR/VR experiences. However, existing methods only generate body motion data, excluding facial expressions and hand movements. This…

Computer Vision and Pattern Recognition · Computer Science 2024-09-23 Mingdian Liu , Yilin Liu , Gurunandan Krishnan , Karl S Bayer , Bing Zhou

We present a method for text-driven perpetual view generation -- synthesizing long-term videos of various scenes solely, given an input text prompt describing the scene and camera poses. We introduce a novel framework that generates such…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Rafail Fridman , Amit Abecasis , Yoni Kasten , Tali Dekel

Text-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Beiyuan Zhang , Yue Ma , Chunlei Fu , Xinyang Song , Zhenan Sun , Ziqiang Li

Diffusion-based methods can generate realistic images and videos, but they struggle to edit existing objects in a video while preserving their appearance over time. This prevents diffusion models from being applied to natural video editing…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Wenhao Chai , Xun Guo , Gaoang Wang , Yan Lu

Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scenes, often collapsing to repetitive layouts, stereotypical poses, and poorly grounded…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Wenxuan Peng , Bharath Hariharan , Hadar Averbuch-Elor

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Renda Li , Xiaohua Qi , Qiang Ling , Jun Yu , Ziyi Chen , Peng Chang , Mei HanJing Xiao

This paper presents a novel framework for speech-driven gesture production, applicable to virtual agents to enhance human-computer interaction. Specifically, we extend recent deep-learning-based, data-driven methods for speech-driven…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Taras Kucherenko , Dai Hasegawa , Naoshi Kaneko , Gustav Eje Henter , Hedvig Kjellström

Text-to-motion generation has advanced rapidly, yet two challenges persist. First, existing motion autoencoders compress each frame into a single monolithic latent vector, entangling trajectory and per-joint rotations in an unstructured…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Zeyu Ling , Qing Shuai , Teng Zhang , Shiyang Li , Bo Han , Changqing Zou

We present a general and simple text to video model based on Transformer. Since both text and video are sequential data, we encode both texts and images into the same hidden space, which are further fed into Transformer to capture the…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Gang Chen

Despite diffusion models having shown powerful abilities to generate photorealistic images, generating videos that are realistic and diverse still remains in its infancy. One of the key reasons is that current methods intertwine spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Zhiwu Qing , Shiwei Zhang , Jiayu Wang , Xiang Wang , Yujie Wei , Yingya Zhang , Changxin Gao , Nong Sang

This paper investigates a novel task of talking face video generation solely from speeches. The speech-to-video generation technique can spark interesting applications in entertainment, customer service, and human-computer-interaction…

Sound · Computer Science 2021-07-15 Shijing Si , Jianzong Wang , Xiaoyang Qu , Ning Cheng , Wenqi Wei , Xinghua Zhu , Jing Xiao

In this study, we introduce T2M-HiFiGPT, a novel conditional generative framework for synthesizing human motion from textual descriptions. This framework is underpinned by a Residual Vector Quantized Variational AutoEncoder (RVQ-VAE) and a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Congyi Wang

Recent techniques for text-to-4D generation synthesize dynamic 3D scenes using supervision from pre-trained text-to-video models. However, existing representations for motion, such as deformation models or time-dependent neural…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Sherwin Bahmani , Xian Liu , Wang Yifan , Ivan Skorokhodov , Victor Rong , Ziwei Liu , Xihui Liu , Jeong Joon Park , Sergey Tulyakov , Gordon Wetzstein , Andrea Tagliasacchi , David B. Lindell

Creation of images using generative adversarial networks has been widely adapted into multi-modal regime with the advent of multi-modal representation models pre-trained on large corpus. Various modalities sharing a common representation…

Sound · Computer Science 2022-06-10 Yoonjeon Kim , Joel Jang , Sumin Shin

We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from…

Computer Vision and Pattern Recognition · Computer Science 2022-09-30 Uriel Singer , Adam Polyak , Thomas Hayes , Xi Yin , Jie An , Songyang Zhang , Qiyuan Hu , Harry Yang , Oron Ashual , Oran Gafni , Devi Parikh , Sonal Gupta , Yaniv Taigman

In recent years, large text-to-video (T2V) synthesis models have garnered considerable attention for their abilities to generate videos from textual descriptions. However, achieving both high imaging quality and effective motion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Tongtong Su , Chengyu Wang , Bingyan Liu , Jun Huang , Dongming Lu
‹ Prev 1 3 4 5 6 7 10 Next ›