中文
相关论文

相关论文: ZeroComp: Zero-shot Object Compositing from Image …

200 篇论文

Zero-shot customized video generation has gained significant attention due to its substantial application potential. Existing methods rely on additional models to extract and inject reference subject features, assuming that the Video…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Tao Wu , Yong Zhang , Xiaodong Cun , Zhongang Qi , Junfu Pu , Huanzhang Dou , Guangcong Zheng , Ying Shan , Xi Li

In recent years, deep neural networks for image inhomogeneity reduction have shown promising results. However, current methods with (un)supervised solutions require preparing a training dataset, which is expensive and laborious for data…

图像与视频处理 · 电气工程与系统科学 2026-02-16 Hongxu Yang , Edina Timko , Brice Fernandez

Gaussian Splatting has become a popular technique for various 3D Computer Vision tasks, including novel view synthesis, scene reconstruction, and dynamic scene rendering. However, the challenge of natural-looking object insertion, where the…

计算机视觉与模式识别 · 计算机科学 2026-05-04 Vsevolod Skorokhodov , Nikita Durasov , Pascal Fua

Large-scale text-to-image diffusion models have significantly improved the state of the art in generative image modelling and allow for an intuitive and powerful user interface to drive the image generation process. Expressing spatial…

计算机视觉与模式识别 · 计算机科学 2023-06-27 Guillaume Couairon , Marlène Careil , Matthieu Cord , Stéphane Lathuilière , Jakob Verbeek

Although diffusion-based zero-shot image restoration and enhancement methods have achieved great success, applying them to video restoration or enhancement will lead to severe temporal flickering. In this paper, we propose the first…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Cong Cao , Huanjing Yue , Shangbin Xie , Xin Liu , Jingyu Yang

We present a learning framework that learns to recover the 3D shape, pose and texture from a single image, trained on an image collection without any ground truth 3D shape, multi-view, camera viewpoints or keypoint supervision. We approach…

计算机视觉与模式识别 · 计算机科学 2020-07-22 Shubham Goel , Angjoo Kanazawa , Jitendra Malik

Recent monocular 3D shape reconstruction methods have shown promising zero-shot results on object-segmented images without any occlusions. However, their effectiveness is significantly compromised in real-world conditions, due to imperfect…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Junhyeong Cho , Kim Youwang , Hunmin Yang , Tae-Hyun Oh

Zero-shot composed image retrieval (ZS-CIR) retrieves a target image from a reference image and a text modification without human-annotated CIR triplets. Projection-based ZS-CIR methods are attractive because they do not rely on LLMs at…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Mingyu Liu , Sihan Huang , Yijia Fan , Yinlin Yan , Quan Zhang , Jian-Fang Hu , Jianhuang Lai

Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive,…

Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part…

计算机视觉与模式识别 · 计算机科学 2021-03-02 Aditya Ramesh , Mikhail Pavlov , Gabriel Goh , Scott Gray , Chelsea Voss , Alec Radford , Mark Chen , Ilya Sutskever

Recently, GAN inversion methods combined with Contrastive Language-Image Pretraining (CLIP) enables zero-shot image manipulation guided by text prompts. However, their applications to diverse real images are still difficult due to the…

计算机视觉与模式识别 · 计算机科学 2022-08-12 Gwanghyun Kim , Taesung Kwon , Jong Chul Ye

Zero-shot object pose estimation enables the retrieval of object poses from images without necessitating object-specific training. In recent approaches this is facilitated by vision foundation models (VFM), which are pre-trained models that…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Bernd Von Gimborn , Philipp Ausserlechner , Markus Vincze , Stefan Thalhammer

Given semantic descriptions of object classes, zero-shot learning aims to accurately recognize objects of the unseen classes, from which no examples are available at the training stage, by associating them to the seen classes, from which…

计算机视觉与模式识别 · 计算机科学 2016-05-31 Soravit Changpinyo , Wei-Lun Chao , Boqing Gong , Fei Sha

Many low-dose CT imaging methods rely on supervised learning, which requires a large number of paired noisy and clean images. However, obtaining paired images in clinical practice is challenging. To address this issue, zero-shot…

图像与视频处理 · 电气工程与系统科学 2025-04-11 Yongyi Shi , Ge Wang

Image composition aims to seamlessly insert a user-specified object into a new scene, but existing models struggle with complex lighting (e.g., accurate shadows, water reflections) and diverse, high-resolution inputs. Modern text-to-image…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Shilin Lu , Zhuming Lian , Zihan Zhou , Shaocong Zhang , Chen Zhao , Adams Wai-Kin Kong

Recent work has shown that diffusion models can serve as powerful neural rendering engines that can be leveraged for inserting virtual objects into images. However, unlike typical physics-based renderers, these neural rendering engines are…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Frédéric Fortier-Chouinard , Zitian Zhang , Louis-Etienne Messier , Mathieu Garon , Anand Bhattad , Jean-François Lalonde

Zero-shot instance segmentation aims to detect and precisely segment objects of unseen categories without any training samples. Since the model is trained on seen categories, there is a strong bias that the model tends to classify all the…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Shuting He , Henghui Ding , Wei Jiang

Zero-shot, training-free, image-based text-to-video generation is an emerging area that aims to generate videos using existing image-based diffusion models. Current methods in this space require specific architectural changes to image…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Diljeet Jagpal , Xi Chen , Vinay P. Namboodiri

Vision-Language Models for remote sensing have shown promising uses thanks to their extensive pretraining. However, their conventional usage in zero-shot scene classification methods still involves dividing large images into patches and…

The success of monocular depth estimation relies on large and diverse training sets. Due to the challenges associated with acquiring dense ground-truth depth across different environments at scale, a number of datasets with distinct…

计算机视觉与模式识别 · 计算机科学 2020-08-26 René Ranftl , Katrin Lasinger , David Hafner , Konrad Schindler , Vladlen Koltun