English
Related papers

Related papers: Motion Guidance: Diffusion-Based Image Editing wit…

200 papers

Image manipulation under the guidance of textual descriptions has recently received a broad range of attention. In this study, we focus on the regional editing of images with the guidance of given text prompts. Different from current…

Computer Vision and Pattern Recognition · Computer Science 2023-02-24 Nisha Huang , Fan Tang , Weiming Dong , Tong-Yee Lee , Changsheng Xu

This paper introduces innovative solutions to enhance spatial controllability in diffusion models reliant on text queries. We first introduce vision guidance as a foundational spatial cue within the perturbed distribution. This…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Zipeng Qi , Guoxi Huang , Chenyang Liu , Fei Ye

Recent advancements in diffusion models have significantly broadened the possibilities for editing images of real-world objects. However, performing non-rigid transformations, such as changing the pose of objects or image-based…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Potito Aghilar , Vito Walter Anelli , Michelantonio Trizio , Tommaso Di Noia

Recent advancements in text-to-image diffusion models have demonstrated remarkable success, yet they often struggle to fully capture the user's intent. Existing approaches using textual inputs combined with bounding boxes or region masks…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Seonho Lee , Jiho Choi , Seohyun Lim , Jiwook Kim , Hyunjung Shim

Recent advancements in Text-to-Image (T2I) diffusion models have demonstrated impressive success in generating high-quality images with zero-shot generalization capabilities. Yet, current models struggle to closely adhere to prompt…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Hyun Kang , Dohae Lee , Myungjin Shin , In-Kwon Lee

Recent advancements in diffusion models have significantly facilitated text-guided video editing. However, there is a relative scarcity of research on image-guided video editing, a method that empowers users to edit videos by merely…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Zhi-Lin Huang , Yixuan Liu , Chujun Qin , Zhongdao Wang , Dong Zhou , Dong Li , Emad Barsoum

Hand pose estimation from a single image has many applications. However, approaches to full 3D body pose estimation are typically trained on day-to-day activities or actions. As such, detailed hand-to-hand interactions are poorly…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Maksym Ivashechkin , Oscar Mendez , Richard Bowden

Pose estimation refers to tracking a human's full body posture, including their head, torso, arms, and legs. The problem is challenging in practical settings where the number of body sensors are limited. Past work has shown promising…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Sahil Bhandary Karnoor , Romit Roy Choudhury

Generative models have enabled intuitive image creation and manipulation using natural language. In particular, diffusion models have recently shown remarkable results for natural image editing. In this work, we propose to apply diffusion…

Large-scale text-to-image diffusion models have significantly improved the state of the art in generative image modelling and allow for an intuitive and powerful user interface to drive the image generation process. Expressing spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Guillaume Couairon , Marlène Careil , Matthieu Cord , Stéphane Lathuilière , Jakob Verbeek

Large-scale generative models, such as text-to-image diffusion models, have garnered widespread attention across diverse domains due to their creative and high-fidelity image generation. Nonetheless, existing large-scale diffusion models…

Computer Vision and Pattern Recognition · Computer Science 2024-08-28 Younghyun Kim , Geunmin Hwang , Junyu Zhang , Eunbyung Park

Diffusion-based image editing models have achieved significant progress in real world applications. However, conventional models typically rely on natural language prompts, which often lack the precision required to localize target objects.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Haohang Xu , Lin Liu , Zhibo Zhang , Rong Cong , Xiaopeng Zhang , Qi Tian

Existing video deraining methods are often trained on paired datasets, either synthetic, which limits their ability to generalize to real-world rain, or captured by static cameras, which restricts their effectiveness in dynamic scenes with…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Tuomas Varanka , Juan Luis Gonzalez , Hyeongwoo Kim , Pablo Garrido , Xu Yao

Image compression technology eliminates redundant information to enable efficient transmission and storage of images, serving both machine vision and human visual perception. For years, image coding focused on human perception has been…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Takahiro Shindo , Yui Tatsumi , Taiju Watanabe , Hiroshi Watanabe

Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks such as optical flow prediction and tracking in a zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Tianqi Zhang , Ziyi Wang , Wenzhao Zheng , Weiliang Chen , Yuanhui Huang , Zhengyang Huang , Jie Zhou , Jiwen Lu

Specifying nuanced and compelling camera motion remains a significant hurdle for non-expert creators using generative tools, creating an "expressive gap" where generic text prompts fail to capture cinematic vision. This barrier limits…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Pooja Guhan , Divya Kothandaraman , Geonsun Lee , Tsung-Wei Huang , Guan-Ming Su , Dinesh Manocha

We propose a method for editing images from human instructions: given an input image and a written instruction that tells the model what to do, our model follows these instructions to edit the image. To obtain training data for this…

Computer Vision and Pattern Recognition · Computer Science 2023-01-19 Tim Brooks , Aleksander Holynski , Alexei A. Efros

Generating high-quality whole-body human object interaction motion sequences is becoming increasingly important in various fields such as animation, VR/AR, and robotics. The main challenge of this task lies in determining the level of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Yonghao Zhang , Qiang He , Yanguang Wan , Yinda Zhang , Xiaoming Deng , Cuixia Ma , Hongan Wang

Generative image editing has recently witnessed extremely fast-paced growth. Some works use high-level conditioning such as text, while others use low-level conditioning. Nevertheless, most of them lack fine-grained control over the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Vidit Goel , Elia Peruzzo , Yifan Jiang , Dejia Xu , Xingqian Xu , Nicu Sebe , Trevor Darrell , Zhangyang Wang , Humphrey Shi

We present Pix2Gif, a motion-guided diffusion model for image-to-GIF (video) generation. We tackle this problem differently by formulating the task as an image translation problem steered by text and motion magnitude prompts, as shown in…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Hitesh Kandala , Jianfeng Gao , Jianwei Yang