English
Related papers

Related papers: MotionBooth: Motion-Aware Customized Text-to-Video…

200 papers

Subject-driven text-to-image generation models create novel renditions of an input subject based on text prompts. Existing models suffer from lengthy fine-tuning and difficulties preserving the subject fidelity. To overcome these…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Dongxu Li , Junnan Li , Steven C. H. Hoi

Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Hila Chefer , Shiran Zada , Roni Paiss , Ariel Ephrat , Omer Tov , Michael Rubinstein , Lior Wolf , Tali Dekel , Tomer Michaeli , Inbar Mosseri

We present Pix2Gif, a motion-guided diffusion model for image-to-GIF (video) generation. We tackle this problem differently by formulating the task as an image translation problem steered by text and motion magnitude prompts, as shown in…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Hitesh Kandala , Jianfeng Gao , Jianwei Yang

Recent advances in diffusion-based text-to-video (T2V) models have demonstrated remarkable progress, but these models still face challenges in generating videos with multiple objects. Most models struggle with accurately capturing complex…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Aimon Rahman , Jiang Liu , Ze Wang , Ximeng Sun , Jialian Wu , Xiaodong Yu , Yusheng Su , Vishal M. Patel , Zicheng Liu , Emad Barsoum

Personalized text-to-image generation methods can generate customized images based on the reference images, which have garnered wide research interest. Recent methods propose a finetuning-free approach with a decoupled cross-attention…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Qihan Huang , Siming Fu , Jinlong Liu , Hao Jiang , Yipeng Yu , Jie Song

Video generation technologies are developing rapidly and have broad potential applications. Among these technologies, camera control is crucial for generating professional-quality videos that accurately meet user expectations. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Wanquan Feng , Jiawei Liu , Pengqi Tu , Tianhao Qi , Mingzhen Sun , Tianxiang Ma , Songtao Zhao , Siyu Zhou , Qian He

Diverse human motion generation is an increasingly important task, having various applications in computer vision, human-computer interaction and animation. While text-to-motion synthesis using diffusion models has shown success in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Heechang Kim , Gwanghyun Kim , Se Young Chun

Multi-object tracking (MOT) has profound applications in a variety of fields, including surveillance, sports analytics, self-driving, and cooperative robotics. Despite considerable advancements, existing MOT methodologies tend to falter…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Hamza Mukhtar , Muhammad Usman Ghani Khan

Accurately distinguishing each object is a fundamental goal of Multi-object tracking (MOT) algorithms. However, achieving this goal still remains challenging, primarily due to: (i) For crowded scenes with occluded objects, the high overlap…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Jiapeng Wu , Yichen Liu

Customized generation has achieved significant progress in image synthesis, yet personalized video generation remains challenging due to temporal inconsistencies and quality degradation. In this paper, we introduce CustomVideoX, an…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 D. She , Mushui Liu , Jingxuan Pang , Jin Wang , Zhen Yang , Wanggui He , Guanghao Zhang , Yi Wang , Qihan Huang , Haobin Tang , Yunlong Yu , Siming Fu

Benefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the parameter-efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Xiuli Bi , Jian Lu , Bo Liu , Xiaodong Cun , Yong Zhang , Weisheng Li , Bin Xiao

In this work, we present Patch-based Object-centric Video Transformer (POVT), a novel region-based video generation architecture that leverages object-centric information to efficiently model temporal dynamics in videos. We build upon prior…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Wilson Yan , Ryo Okumura , Stephen James , Pieter Abbeel

We have recently seen tremendous progress in realistic text-to-motion generation. Yet, the existing methods often fail or produce implausible motions with unseen text inputs, which limits the applications. In this paper, we present OMG, a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Han Liang , Jiacheng Bao , Ruichi Zhang , Sihan Ren , Yuecheng Xu , Sibei Yang , Xin Chen , Jingyi Yu , Lan Xu

Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. However, current…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Ruiyan Wang , Teng Hu , Kaihui Huang , Zihan Su , Ran Yi , Lizhuang Ma

Text-driven motion generation offers a powerful and intuitive way to create human movements directly from natural language. By removing the need for predefined motion inputs, it provides a flexible and accessible approach to controlling…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Ali Rida Sahili , Najett Neji , Hedi Tabia

Whole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to achieve various generation tasks with…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Yuxuan Bian , Ailing Zeng , Xuan Ju , Xian Liu , Zhaoyang Zhang , Wei Liu , Qiang Xu

We introduce precise object silhouette as a new form of user control in text-to-image diffusion models, which we dub Shape-Guided Diffusion. Our training-free method uses an Inside-Outside Attention mechanism during the inversion and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Dong Huk Park , Grace Luo , Clayton Toste , Samaneh Azadi , Xihui Liu , Maka Karalashvili , Anna Rohrbach , Trevor Darrell

Text-based 3D human motion editing is a critical yet challenging task in computer vision and graphics. While training-free approaches have been explored, the recent release of the MotionFix dataset, which includes source-text-motion…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Zhengyuan Li , Kai Cheng , Anindita Ghosh , Uttaran Bhattacharya , Liangyan Gui , Aniket Bera

Recently, diffusion models like StableDiffusion have achieved impressive image generation results. However, the generation process of such diffusion models is uncontrollable, which makes it hard to generate videos with continuous and…

Computer Vision and Pattern Recognition · Computer Science 2023-08-04 Zhihao Hu , Dong Xu

Diffusion models have demonstrated impressive image generation capabilities. Personalized approaches, such as textual inversion and Dreambooth, enhance model individualization using specific images. These methods enable generating images of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Yan Zeng , Masanori Suganuma , Takayuki Okatani