中文
相关论文

相关论文: SteeredMarigold: Steering Diffusion Towards Depth …

200 篇论文

Planning with pretrained diffusion models has emerged as a promising approach for solving test-time guided control problems. Standard gradient guidance typically performs optimally under convex, differentiable reward landscapes. However, it…

人工智能 · 计算机科学 2025-11-11 Hyeonseong Jeon , Cheolhong Min , Jaesik Park

The secure analysis of dermatological images in clinical environments is fundamentally restricted by the critical trade-off between patient privacy and the preservation of diagnostic fidelity. Traditional de-identification techniques often…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Konstantinos Moutselos , Ilias Maglogiannis

3D reconstruction from a single image is a long-standing problem in computer vision. Learning-based methods address its inherent scale ambiguity by leveraging increasingly large labeled and unlabeled datasets, to produce geometric priors…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Vitor Guizilini , Pavel Tokmakov , Achal Dave , Rares Ambrus

Hiding data using neural networks (i.e., neural steganography) has achieved remarkable success across both discriminative classifiers and generative adversarial networks. However, the potential of data hiding in diffusion models remains…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Haoyu Chen , Yunqiao Yang , Nan Zhong , Kede Ma

Depth completion involves predicting dense depth maps from sparse LiDAR inputs. However, sparse depth annotations from sensors limit the availability of dense supervision, which is necessary for learning detailed geometric features. In this…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Yingping Liang , Yutao Hu , Wenqi Shao , Ying Fu

Monocular camera calibration is a key precondition for numerous 3D vision applications. Despite considerable advancements, existing methods often hinge on specific assumptions and struggle to generalize across varied real-world scenarios,…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Xiankang He , Guangkai Xu , Bo Zhang , Hao Chen , Ying Cui , Dongyan Guo

Accurate completion and denoising of roof height maps are crucial to reconstructing high-quality 3D buildings. Repairing sparse points can enhance low-cost sensor use and reduce UAV flight overlap. RoofDiffusion is a new end-to-end…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Kyle Shih-Huang Lo , Jörg Peters , Eric Spellman

Recent score-based diffusion models (SBDMs) show promising results in unpaired image-to-image translation (I2I). However, existing methods, either energy-based or statistically-based, provide no explicit form of the interfered intermediate…

计算机视觉与模式识别 · 计算机科学 2023-08-07 Shikun Sun , Longhui Wei , Junliang Xing , Jia Jia , Qi Tian

Recent advances in diffusion models have demonstrated their strong capabilities in generating high-fidelity samples from complex distributions through an iterative refinement process. Despite the empirical success of diffusion models in…

机器人学 · 计算机科学 2024-07-03 Chaoyi Pan , Zeji Yi , Guanya Shi , Guannan Qu

Large-scale text-to-image diffusion models have significantly improved the state of the art in generative image modelling and allow for an intuitive and powerful user interface to drive the image generation process. Expressing spatial…

计算机视觉与模式识别 · 计算机科学 2023-06-27 Guillaume Couairon , Marlène Careil , Matthieu Cord , Stéphane Lathuilière , Jakob Verbeek

Active depth cameras suffer from several limitations, which cause incomplete and noisy depth maps, and may consequently affect the performance of RGB-D Odometry. To address this issue, this paper presents a visual odometry method based on…

机器人学 · 计算机科学 2017-08-10 Pedro F. Proença , Yang Gao

Due to the visual properties of reflection and refraction, RGB-D cameras cannot accurately capture the depth of transparent objects, leading to incomplete depth maps. To fill in the missing points, recent studies tend to explore new visual…

计算机视觉与模式识别 · 计算机科学 2024-08-02 Yiheng Huang , Junhong Chen , Nick Michiels , Muhammad Asim , Luc Claesen , Wenyin Liu

Image Auto-regressive (AR) models have emerged as a powerful paradigm of visual generative models. Despite their promising performance, they suffer from slow generation speed due to the large number of sampling steps required. Although…

机器学习 · 计算机科学 2025-10-27 Enshu Liu , Qian Chen , Xuefei Ning , Shengen Yan , Guohao Dai , Zinan Lin , Yu Wang

Recent video depth estimation methods achieve great performance by following the paradigm of image depth estimation, i.e., typically fine-tuning pre-trained video diffusion models with massive data. However, we argue that video depth…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Haodong Li , Chen Wang , Jiahui Lei , Kostas Daniilidis , Lingjie Liu

In this paper, we present the Directly Denoising Diffusion Model (DDDM): a simple and generic approach for generating realistic images with few-step sampling, while multistep sampling is still preserved for better performance. DDDMs require…

计算机视觉与模式识别 · 计算机科学 2024-06-03 Dan Zhang , Jingjing Wang , Feng Luo

In this work, we present a panoramic metric depth foundation model that generalizes across diverse scene distances. We explore a data-in-the-loop paradigm from the view of both data construction and framework design. We collect a…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Xin Lin , Meixi Song , Dizhe Zhang , Wenxuan Lu , Haodong Li , Bo Du , Ming-Hsuan Yang , Truong Nguyen , Lu Qi

Machine learning algorithms have enabled high quality stereo depth estimation to run on Augmented and Virtual Reality (AR/VR) devices. However, high energy consumption across the full image processing stack prevents stereo depth algorithms…

计算机视觉与模式识别 · 计算机科学 2025-02-14 Jack Erhardt , Ziang Li , Reid Pinkham , Andrew Berkovich , Zhengya Zhang

Generative steganography (GS) is an emerging technique that generates stego images directly from secret data. Various GS methods based on GANs or Flow have been developed recently. However, existing GAN-based GS methods cannot completely…

多媒体 · 计算机科学 2023-09-07 Ping Wei , Qing Zhou , Zichi Wang , Zhenxing Qian , Xinpeng Zhang , Sheng Li

Recent advances in zero-shot text-to-3D human generation, which employ the human model prior (eg, SMPL) or Score Distillation Sampling (SDS) with pre-trained text-to-image diffusion models, have been groundbreaking. However, SDS may provide…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Jianhui Yu , Hao Zhu , Liming Jiang , Chen Change Loy , Weidong Cai , Wayne Wu

Diffusion-based text-to-image generation models trained on extensive text-image pairs have demonstrated the ability to produce photorealistic images aligned with textual descriptions. However, a significant limitation of these models is…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Mingyuan Zhou , Zhendong Wang , Huangjie Zheng , Hai Huang