English
Related papers

Related papers: Diffusion Features for Zero-Shot 6DoF Object Pose …

200 papers

Image diffusion models, though originally developed for image generation, implicitly capture rich semantic structures that enable various recognition and localization tasks beyond synthesis. In this work, we investigate their self-attention…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Youngseo Kim , Dohyun Kim , Geonhee Han , Paul Hongsuck Seo

We propose a diffusion-based framework for zero-shot image editing that unifies text-guided and reference-guided approaches without requiring fine-tuning. Our method leverages diffusion inversion and timestep-specific null-text embeddings…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Dasol Jeong , Donggoo Kang , Jiwon Park , Hyebean Lee , Joonki Paik

Deep learning underlies most modern approaches and tools in computer vision, including biomedical imaging. However, for interactive semantic segmentation (often called pixel classification in this context) and interactive object-level…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Carolin Teuber , Anwai Archit , Tobias Boothe , Peter Ditte , Jochen Rink , Constantin Pape

Low-field (LF) magnetic resonance imaging (MRI) democratizes access to diagnostic imaging but is fundamentally limited by low signal-to-noise ratio and significant tissue contrast distortion due to field-dependent relaxation dynamics.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Muyu Liu , Chenhe Du , Xuanyu Tian , Qing Wu , Xiao Wang , Haonan Zhang , Hongjiang Wei , Yuyao Zhang

While object reconstruction has made great strides in recent years, current methods typically require densely captured images and/or known camera poses, and generalize poorly to novel object categories. To step toward object reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2024-01-29 Hanwen Jiang , Zhenyu Jiang , Kristen Grauman , Yuke Zhu

Despite the significant progress in six degrees-of-freedom (6DoF) object pose estimation, existing methods have limited applicability in real-world scenarios involving embodied agents and downstream 3D vision tasks. These limitations mainly…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Zhiwen Fan , Panwang Pan , Peihao Wang , Yifan Jiang , Dejia Xu , Hanwen Jiang , Zhangyang Wang

Existing works on video frame interpolation (VFI) mostly employ deep neural networks that are trained by minimizing the L1, L2, or deep feature space distance (e.g. VGG loss) between their outputs and ground-truth frames. However, recent…

Image and Video Processing · Electrical Eng. & Systems 2024-06-11 Duolikun Danier , Fan Zhang , David Bull

One practical approach to infer 3D scene structure from a single image is to retrieve a closely matching 3D model from a database and align it with the object in the image. Existing methods rely on supervised training with images and pose…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Pattaramanee Arsomngern , Sasikarn Khwanmuang , Matthias Nießner , Supasorn Suwajanakorn

Recent advances in large generative models have shown that simple autoregressive formulations, when scaled appropriately, can exhibit strong zero-shot generalization across domains. Motivated by this trend, we investigate whether…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Yuxiang Lai , Jike Zhong , Ming Li , Yuheng Li , Xiaofeng Yang

Typical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the correct template often leads to inaccurate pose…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Junwen Huang , Shishir Reddy Vutukur , Peter KT Yu , Nassir Navab , Slobodan Ilic , Benjamin Busam

Denoising diffusion probabilistic models have transformed image generation with their impressive fidelity and diversity. We show that they also excel in estimating optical flow and monocular depth, surprisingly, without task-specific…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Saurabh Saxena , Charles Herrmann , Junhwa Hur , Abhishek Kar , Mohammad Norouzi , Deqing Sun , David J. Fleet

Given sparse views of a 3D object, estimating their camera poses is a long-standing and intractable problem. Toward this goal, we consider harnessing the pre-trained diffusion model of novel views conditioned on viewpoints (Zero-1-to-3). We…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Weihao Cheng , Yan-Pei Cao , Ying Shan

Human has an incredible ability to effortlessly perceive the viewpoint difference between two images containing the same object, even when the viewpoint change is astonishingly vast with no co-visible regions in the images. This remarkable…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Yujing Sun , Caiyi Sun , Yuan Liu , Yuexin Ma , Siu Ming Yiu

Behavioral Foundation Models (BFMs) proved successful in producing policies for arbitrary tasks in a zero-shot manner, requiring no test-time training or task-specific fine-tuning. Among the most promising BFMs are the ones that estimate…

Machine Learning · Computer Science 2026-05-05 Maksim Bobrin , Ilya Zisman , Alexander Nikulin , Vladislav Kurenkov , Dmitry Dylov

We propose FoundPose, a model-based method for 6D pose estimation of unseen objects from a single RGB image. The method can quickly onboard new objects using their 3D models without requiring any object- or task-specific training. In…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Evin Pınar Örnek , Yann Labbé , Bugra Tekin , Lingni Ma , Cem Keskin , Christian Forster , Tomas Hodan

Large-scale text-to-image diffusion models have shown impressive capabilities for generative tasks by leveraging strong vision-language alignment from pre-training. However, most vision-language discriminative tasks require extensive…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Xuyang Liu , Siteng Huang , Yachen Kang , Honggang Chen , Donglin Wang

6D pose estimation is a central problem in robot vision. Compared with pose estimation based on point correspondences or its robust versions, correspondence-free methods are often more flexible. However, existing correspondence-free methods…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Quan Quan , Dun Dai

Recently, the diffusion model has emerged as a superior generative model that can produce high quality and realistic images. However, for medical image translation, the existing diffusion models are deficient in accurately retaining…

Image and Video Processing · Electrical Eng. & Systems 2023-10-31 Yunxiang Li , Hua-Chieh Shao , Xiao Liang , Liyuan Chen , Ruiqi Li , Steve Jiang , Jing Wang , You Zhang

Estimating the 6D pose of objects from RGBD data is a fundamental problem in computer vision, with applications in robotics and augmented reality. A key challenge is achieving generalization to novel objects that were not seen during…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Andrea Caraffa , Davide Boscaini , Fabio Poiesi

Estimating relative camera poses between images has been a central problem in computer vision. Methods that find correspondences and solve for the fundamental matrix offer high precision in most cases. Conversely, methods predicting pose…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Chris Rockwell , Nilesh Kulkarni , Linyi Jin , Jeong Joon Park , Justin Johnson , David F. Fouhey