中文
相关论文

相关论文: FreeInv: Free Lunch for Improving DDIM Inversion

200 篇论文

Synthesizing novel views from a single input image is a challenging task. It requires extrapolating the 3D structure of a scene while inferring details in occluded regions, and maintaining geometric consistency across viewpoints. Many…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Sehajdeep Singh , A V Subramanyam , Aditya Gupta , Sahil Gupta

In this study, we aim to determine and solve the deficiency of Stable Diffusion Inpainting (SDI) in following the instruction of both prompt and mask. Due to the training bias from masking, the inpainting quality is hindered when the prompt…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Teng-Fang Hsiao , Bo-Kai Ruan , Sung-Lin Tsai , Yi-Lun Wu , Hong-Han Shuai

Diffusion Transformer (DiT), an emerging diffusion model for image generation, has demonstrated superior performance but suffers from substantial computational costs. Our investigations reveal that these costs stem from the static inference…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Wangbo Zhao , Yizeng Han , Jiasheng Tang , Kai Wang , Yibing Song , Gao Huang , Fan Wang , Yang You

There is a growing interest in the use of latent diffusion models (LDMs) for image restoration (IR) tasks due to their ability to model effectively the distribution of natural images. While significant progress has been made, there are…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Di You , Daniel Siromani , Pier Luigi Dragotti

Creating 3D assets from single-view images is a complex task that demands a deep understanding of the world. Recently, feed-forward 3D generative models have made significant progress by training large reconstruction models on extensive 3D…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Wenqiang Sun , Zhengyi Wang , Shuo Chen , Yikai Wang , Zilong Chen , Jun Zhu , Jun Zhang

Computational imaging methods increasingly rely on powerful generative diffusion models to tackle challenging image restoration tasks. In particular, state-of-the-art zero-shot image inverse solvers leverage distilled text-to-image latent…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Alessio Spagnoletti , Andrés Almansa , Marcelo Pereyra

Rectified-Flow (RF)-based generative models have recently emerged as strong alternatives to traditional diffusion models, demonstrating state-of-the-art performance across various tasks. By learning a continuous velocity field that…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Chenru Wang , Beier Zhu , Chi Zhang

Recent advances in inverse problem solving have increasingly adopted flow priors over diffusion models due to their ability to construct straight probability paths from noise to data, thereby enhancing efficiency in both training and…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Hossein Askari , Yadan Luo , Hongfu Sun , Fred Roosta

Deep image prior (DIP) is a recently proposed technique for solving imaging inverse problems by fitting the reconstructed images to the output of an untrained convolutional neural network. Unlike pretrained feedforward neural networks, the…

计算机视觉与模式识别 · 计算机科学 2022-09-20 Kevin Zhang , Mingyang Xie , Maharshi Gor , Yi-Ting Chen , Yvonne Zhou , Christopher A. Metzler

Training-free image editing has attracted increasing attention for its efficiency and independence from training data. However, existing approaches predominantly rely on inversion-reconstruction trajectories, which impose an inherent…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Menglin Han , Zhangkai Ni

Diffusion Transformer (DiT), an emerging diffusion model for visual generation, has demonstrated superior performance but suffers from substantial computational costs. Our investigations reveal that these costs primarily stem from the…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Wangbo Zhao , Yizeng Han , Jiasheng Tang , Kai Wang , Hao Luo , Yibing Song , Gao Huang , Fan Wang , Yang You

Multi-view inverse rendering aims to recover geometry, materials, and illumination consistently across multiple viewpoints. When applied to multi-view images, existing single-view approaches often ignore cross-view relationships, leading to…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Xiangzuo Wu , Chengwei Ren , Jun Zhou , Xiu Li , Yuan Liu

Diffusion models have shown to be strong representation learners, showcasing state-of-the-art performance across multiple domains. Aside from accelerated sampling, DDIM also enables the inversion of real images back to their latent codes. A…

人工智能 · 计算机科学 2025-10-02 Seunghoo Hong , Geonho Son , Juhun Lee , Simon S. Woo

Large-scale vision foundation models such as DINOv2 boast impressive performances by leveraging massive architectures and training datasets. But numerous scenarios require practitioners to reproduce those pre-training solutions, such as on…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Jiaqi Zhang , Juntuo Wang , Zhixin Sun , John Zou , Randall Balestriero

Accurate trajectory prediction is a cornerstone for the safe operation of autonomous driving systems, where understanding the dynamic behavior of surrounding agents is crucial. Transformer-based architectures have demonstrated significant…

机器学习 · 计算机科学 2025-05-07 JianLin Zhu , HongKuo Niu

Despite all recent progress, it is still challenging to edit and manipulate natural images with modern generative models. When using Generative Adversarial Network (GAN), one major hurdle is in the inversion process mapping a real image to…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Zhihong Pan , Riccardo Gherardi , Xiufeng Xie , Stephen Huang

Diffusion models deliver high-fidelity synthesis but remain slow due to iterative sampling. We empirically observe there exists feature invariance in deterministic sampling, and present InvarDiff, a training-free acceleration method that…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Zihao Wu

Diffusion models have emerged as powerful tools for solving inverse problems due to their exceptional ability to model complex prior distributions. However, existing methods predominantly assume known forward operators (i.e., non-blind),…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Weimin Bai , Siyi Chen , Wenzheng Chen , He Sun

Novel view synthesis is required in many robotic applications, such as VR teleoperation and scene reconstruction. Existing methods are often too slow for these contexts, cannot handle dynamic scenes, and are limited by their explicit depth…

计算机视觉与模式识别 · 计算机科学 2022-05-16 Andre Rochow , Max Schwarz , Michael Weinmann , Sven Behnke

Diffusion models have made tremendous progress in text-driven image and video generation. Now text-to-image foundation models are widely applied to various downstream image synthesis tasks, such as controllable image generation and image…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Fengyuan Shi , Jiaxi Gu , Hang Xu , Songcen Xu , Wei Zhang , Limin Wang