English
Related papers

Related papers: Plan, Posture and Go: Towards Open-World Text-to-M…

200 papers

We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Jing Tan , Zhaoyang Zhang , Yantao Shen , Jiarui Cai , Shuo Yang , Jiajun Wu , Wei Xia , Zhuowen Tu , Stefano Soatto

Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mengchao Wang , Qiang Wang , Fan Jiang , Yaqi Fan , Yunpeng Zhang , Yonggang Qi , Kun Zhao , Mu Xu

Denoising diffusion models have shown great promise in human motion synthesis conditioned on natural language descriptions. However, integrating spatial constraints, such as pre-defined motion trajectories and obstacles, remains a challenge…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Korrawe Karunratanakul , Konpat Preechakul , Supasorn Suwajanakorn , Siyu Tang

Text-to-image diffusion models have demonstrated an impressive ability to produce high-quality outputs. However, they often struggle to accurately follow fine-grained spatial information in an input text. To this end, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Ran Galun , Sagie Benaim

Text-driven human motion generation is an emerging task in animation and humanoid robot design. Existing algorithms directly generate the full sequence which is computationally expensive and prone to errors as it does not pay special…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Zichen Geng , Caren Han , Zeeshan Hayder , Jian Liu , Mubarak Shah , Ajmal Mian

This paper addresses the challenge of text-conditioned streaming motion generation, which requires us to predict the next-step human pose based on variable-length historical motions and incoming texts. Existing methods struggle to achieve…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Lixing Xiao , Shunlin Lu , Huaijin Pi , Ke Fan , Liang Pan , Yueer Zhou , Ziyong Feng , Xiaowei Zhou , Sida Peng , Jingbo Wang

We develop a diffusion-based approach for various document layout sequence generation. Layout sequences specify the contents of a document design in an explicit format. Our novel diffusion-based approach works in the sequence domain rather…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Liu He , Yijuan Lu , John Corring , Dinei Florencio , Cha Zhang

Language-image pre-training is an effective technique for learning powerful representations in general domains. However, when directly turning to person representation learning, these general pre-training methods suffer from unsatisfactory…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Jialong Zuo , Jiahao Hong , Feng Zhang , Changqian Yu , Hanyu Zhou , Changxin Gao , Nong Sang , Jingdong Wang

Diffusion-based generative models have significantly advanced text-to-image generation but encounter challenges when processing lengthy and intricate text prompts describing complex scenes with multiple objects. While excelling in…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Hanan Gani , Shariq Farooq Bhat , Muzammal Naseer , Salman Khan , Peter Wonka

This paper targets to enhance the diffusion-based text-to-video generation by improving the two input prompts, including the noise and the text. Accommodated with this goal, we propose POS, a training-free Prompt Optimization Suite to boost…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Shijie Ma , Huayi Xu , Mengjian Li , Weidong Geng , Yaxiong Wang , Meng Wang

Generating realistic robotic manipulation videos is an important step toward unifying perception, planning, and action in embodied agents. While existing video diffusion models require large domain-specific datasets and struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Ye Pang

Generating 3D human motions from textual descriptions is an important research problem with broad applications in video games, virtual reality, and augmented reality. Recent methods align the textual description with human motion at the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Bowen Dang , Lin Wu , Xiaohang Yang , Zheng Yuan , Zhixiang Chen

In the last two years, text-to-image diffusion models have become extremely popular. As their quality and usage increase, a major concern has been the need for better output control. In addition to prompt engineering, one effective method…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Clément Bonnet , Ariel N. Lee , Franck Wertel , Antoine Tamano , Tanguy Cizain , Pablo Ducru

We introduce the Cross Human Motion Diffusion Model (CrossDiff), a novel approach for generating high-quality human motion based on textual descriptions. Our method integrates 3D and 2D information using a shared transformer network within…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Zeping Ren , Shaoli Huang , Xiu Li

Significant advancements in video diffusion models have brought substantial progress to the field of text-to-video (T2V) synthesis. However, existing T2V synthesis model struggle to accurately generate complex motion dynamics, leading to a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Haoran Cheng , Liang Peng , Linxuan Xia , Yuepeng Hu , Hengjia Li , Qinglin Lu , Xiaofei He , Boxi Wu

Generative models have enabled intuitive image creation and manipulation using natural language. In particular, diffusion models have recently shown remarkable results for natural image editing. In this work, we propose to apply diffusion…

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains challenging and untouched so far. Fortunately, human motion…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Biao Jiang , Xin Chen , Wen Liu , Jingyi Yu , Gang Yu , Tao Chen

Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets and paired text descriptions. However, how to effectively…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Xiaoyan Cong , Zekun Li , Zhiyang Dou , Hongyu Li , Omid Taheri , Chuan Guo , Abhay Mittal , Sizhe An , Taku Komura , Wojciech Matusik , Michael J. Black , Srinath Sridhar

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, while generating…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Xingyu Chen

Text-to-motion (T2M) generation is becoming a practical tool for animation and interactive avatars. However, modifying specific body parts while maintaining overall motion coherence remains challenging. Existing methods typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Minyue Dai , Ke Fan , Anyi Rao , Jingbo Wang , Bo Dai