English
Related papers

Related papers: Enhancing Motion in Text-to-Video Generation with …

200 papers

Text-to-motion (T2M) generation has broad applications in character animation, virtual avatars, and human-robot interaction. Existing methods typically generate pose trajectories or motion tokens directly from language, forcing a single…

Machine Learning · Computer Science 2026-05-29 Nikolay Shvetsov , Maksim Bobrin , Nazar Buzun , Dmitry V. Dylov

Video captioning is an advanced multi-modal task which aims to describe a video clip using a natural language sentence. The encoder-decoder framework is the most popular paradigm for this task in recent years. However, there exist some…

Computer Vision and Pattern Recognition · Computer Science 2021-02-15 Haoran Chen , Jianmin Li , Xiaolin Hu

We introduce an approach to generating videos based on a series of given language descriptions. Frames of the video are generated sequentially and optimized by guidance from the CLIP image-text encoder; iterating through language…

Computer Vision and Pattern Recognition · Computer Science 2022-10-25 Peter Schaldenbrand , Zhixuan Liu , Jean Oh

In this paper, we address the challenging problem of long-term 3D human motion generation. Specifically, we aim to generate a long sequence of smoothly connected actions from a stream of multiple sentences (i.e., paragraph). Previous…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Taeryung Lee , Fabien Baradel , Thomas Lucas , Kyoung Mu Lee , Gregory Rogez

We address the problem of generating diverse 3D human motions from textual descriptions. This challenging task requires joint modeling of both modalities: understanding and extracting useful human-centric information from the text, and then…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Mathis Petrovich , Michael J. Black , Gül Varol

Text-to-image (T2I) diffusion models have revolutionized visual content creation, but extending these capabilities to text-to-video (T2V) generation remains a challenge, particularly in preserving temporal consistency. Existing methods that…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Dohun Lee , Bryan S Kim , Geon Yeong Park , Jong Chul Ye

Text-driven motion generation offers a powerful and intuitive way to create human movements directly from natural language. By removing the need for predefined motion inputs, it provides a flexible and accessible approach to controlling…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Ali Rida Sahili , Najett Neji , Hedi Tabia

Dexterous manipulation remains a challenging robotics problem, largely due to the difficulty of collecting extensive human demonstrations for learning. In this paper, we introduce \textsc{Gen2Real}, which replaces costly human demos with…

Robotics · Computer Science 2025-09-18 Kai Ye , Yuhang Wu , Shuyuan Hu , Junliang Li , Meng Liu , Yongquan Chen , Rui Huang

Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Shuling Zhao , Fa-Ting Hong , Xiaoshui Huang , Dan Xu

Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dynamic visual information and static semantics, thereby…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Dang Jisheng , Wu Xudong , Wang Bimei , Lv Ning , Chen Jiayu , Jingwen Zhao , Yichu liu , Jizhao Liu , Juncheng Li , Teng Wang

We introduce MoLingo, a text-to-motion (T2M) model that generates realistic, lifelike human motion by denoising in a continuous latent space. Recent works perform latent space diffusion, either on the whole latent at once or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yannan He , Garvita Tiwari , Xiaohan Zhang , Pankaj Bora , Tolga Birdal , Jan Eric Lenssen , Gerard Pons-Moll

Recent audio-video generative systems suggest that coupling modalities benefits not only audio-video synchrony but also the video modality itself. We pose a fundamental question: Does audio-video joint denoising training improve video…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Jianzong Wu , Hao Lian , Dachao Hao , Ye Tian , Qingyu Shi , Biaolong Chen , Hao Jiang , Yunhai Tong

We present T2Bs, a framework for generating high-quality, animatable character head morphable models from text by combining static text-to-3D generation with video diffusion. Text-to-3D models produce detailed static geometry but lack…

Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yifang Men , Yuan Yao , Miaomiao Cui , Liefeng Bo

Motion synthesis in real-world 3D scenes has recently attracted much attention. However, the static environment assumption made by most current methods usually cannot be satisfied especially for real-time motion synthesis in scanned point…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Jingyu Gong , Min Wang , Wentao Liu , Chen Qian , Zhizhong Zhang , Yuan Xie , Lizhuang Ma

Recent advancements in portrait video generation have been noteworthy. However, existing methods rely heavily on human priors and pre-trained generative models, Motion representations based on human priors may introduce unrealistic motion,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Qiyuan Zhang , Chenyu Wu , Wenzhang Sun , Huaize Liu , Donglin Di , Wei Chen , Changqing Zou

Generative AI has made significant strides in computer vision, particularly in text-driven image/video synthesis (T2I/T2V). Despite the notable advancements, it remains challenging in human-centric content synthesis such as realistic dance…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Tan Wang , Linjie Li , Kevin Lin , Yuanhao Zhai , Chung-Ching Lin , Zhengyuan Yang , Hanwang Zhang , Zicheng Liu , Lijuan Wang

Enabling humanoid robots to synthesize complex, physically coherent motions from natural language commands is a cornerstone of autonomous robotics and human-robot interaction. While diffusion models have shown promise in this text-to-motion…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Wenshuo Chen , Haozhe Jia , Songning Lai , Lei Wang , Yuqi Lin , Hongru Xiao , Lijie Hu , Yutao Yue

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Jiasong Feng , Ao Ma , Jing Wang , Ke Cao , Zhanjie Zhang

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Ashkan Taghipour , Morteza Ghahremani , Zinuo Li , Hamid Laga , Farid Boussaid , Mohammed Bennamoun
‹ Prev 1 4 5 6 7 8 10 Next ›