English
Related papers

Related papers: Zero-to-Hero: Zero-Shot Initialization Empowering …

200 papers

Diffusion model-based low-light image enhancement methods rely heavily on paired training data, leading to limited extensive application. Meanwhile, existing unsupervised methods lack effective bridging capabilities for unknown degradation.…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Jinhong He , Minglong Xue , Aoxiang Ning , Chengyun Song

The current dominant paradigm for imitation learning relies on strong supervision of expert actions to learn both 'what' and 'how' to imitate. We pursue an alternative paradigm wherein an agent first explores the world without any expert…

Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterpart, motion…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Yixuan Ren , Yang Zhou , Jimei Yang , Jing Shi , Difan Liu , Feng Liu , Mingi Kwon , Abhinav Shrivastava

Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part…

Computer Vision and Pattern Recognition · Computer Science 2021-03-02 Aditya Ramesh , Mikhail Pavlov , Gabriel Goh , Scott Gray , Chelsea Voss , Alec Radford , Mark Chen , Ilya Sutskever

Recent advances in training-free attention control methods have enabled flexible and efficient text-guided editing capabilities for existing generation models. However, current approaches struggle to simultaneously deliver strong editing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Zixin Yin , Ling-Hao Chen , Lionel Ni , Xili Dai

Zero-Shot Action Recognition has attracted attention in the last years and many approaches have been proposed for recognition of objects, events and actions in images and videos. There is a demand for methods that can classify instances…

Computer Vision and Pattern Recognition · Computer Science 2021-12-17 Valter Estevam , Helio Pedrini , David Menotti

Recent developments in generative diffusion models have turned many dreams into realities. For video object insertion, existing methods typically require additional information, such as a reference video or a 3D asset of the object, to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Qi Zhao , Zhan Ma , Pan Zhou

We propose \textbf{IC-Effect}, an instruction-guided, DiT-based framework for few-shot video VFX editing that synthesizes complex effects (\eg flames, particles and cartoon characters) while strictly preserving spatial and temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Yuanhang Li , Yiren Song , Junzhe Bai , Xinran Liang , Hu Yang , Libiao Jin , Qi Mao

The optimization of viewers' quality of experience (QoE) in 360 videos faces two major roadblocks: inaccurate adaptive streaming and viewers missing the plot of a story. Alignment edit emerged as a promising mechanism to avoid both issues…

Multimedia · Computer Science 2022-05-24 Lucas Althoff , Alessandro Rodrigues , Mylène C. Q. Farias

High-fidelity surgical video generation can greatly improve medical training and the development of AI, adapting these generative models for precise video editing remains a formidable challenge. Modifying surgical attributes, such as…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Ritul Jangir , Arkya Jyoti Bagchi , Aiman Farooq , Mangalton Okram , Saurabh Seetaram Korgaonkar , Deepak Mishra

Zero-shot recognition aims to classify an image by selecting the most compatible label description from a set of candidate classes without any task-specific supervision. In fine-grained settings, however, the relevant evidence often lies in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Junyi Hu , Qiji Zhou , Lei Zhang , Yue Zhang

Although image editing techniques have advanced significantly, video editing, which aims to manipulate videos according to user intent, remains an emerging challenge. Most existing image-conditioned video editing methods either require…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Xianghao Kong , Hansheng Chen , Yuwei Guo , Lvmin Zhang , Gordon Wetzstein , Maneesh Agrawala , Anyi Rao

Zero-shot Text-to-Video synthesis generates videos based on prompts without any videos. Without motion information from videos, motion priors implied in prompts are vital guidance. For example, the prompt "airplane landing on the runway"…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Sitong Su , Litao Guo , Lianli Gao , Hengtao Shen , Jingkuan Song

Image restoration has traditionally required training specialized models on thousands of paired examples per degradation type. We challenge this paradigm by demonstrating that powerful pre-trained text-conditioned image editing models can…

Image and Video Processing · Electrical Eng. & Systems 2026-01-21 M. Akın Yılmaz , Ahmet Bilican , Burak Can Biner , A. Murat Tekalp

High-resolution content creation is rapidly emerging as a central challenge in both the vision and graphics communities. Images serve as the most fundamental modality for visual expression, and content generation that aligns with the user…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Junsung Lee , Hyunsoo Lee , Yong Jae Lee , Bohyung Han

Humans can easily segment moving objects without knowing what they are. That objectness could emerge from continuous visual observations motivates us to model grouping and movement concurrently from unlabeled videos. Our premise is that a…

Computer Vision and Pattern Recognition · Computer Science 2021-11-12 Runtao Liu , Zhirong Wu , Stella X. Yu , Stephen Lin

Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., "a woman is drinking water."). Existing TI2V frameworks often require…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Haomiao Ni , Bernhard Egger , Suhas Lohit , Anoop Cherian , Ye Wang , Toshiaki Koike-Akino , Sharon X. Huang , Tim K. Marks

Reference-to-video (R2V) generation aims to synthesize videos that align with a text prompt while preserving the subject identity from reference images. However, current R2V methods are hindered by the reliance on explicit reference…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Zijian Zhou , Shikun Liu , Haozhe Liu , Haonan Qiu , Zhaochong An , Weiming Ren , Zhiheng Liu , Xiaoke Huang , Kam Woh Ng , Tian Xie , Xiao Han , Yuren Cong , Hang Li , Chuyan Zhu , Aditya Patel , Tao Xiang , Sen He

Generating free-viewpoint videos is critical for immersive VR/AR experience but recent neural advances still lack the editing ability to manipulate the visual perception for large dynamic scenes. To fill this gap, in this paper we propose…

Computer Vision and Pattern Recognition · Computer Science 2021-05-03 Jiakai Zhang , Xinhang Liu , Xinyi Ye , Fuqiang Zhao , Yanshun Zhang , Minye Wu , Yingliang Zhang , Lan Xu , Jingyi Yu

We introduce an object-aware decoder for improving the performance of spatio-temporal representations on ego-centric videos. The key idea is to enhance object-awareness during training by tasking the model to predict hand positions, object…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Chuhan Zhang , Ankush Gupta , Andrew Zisserman
‹ Prev 1 3 4 5 6 7 10 Next ›