English
Related papers

Related papers: PanDA: Towards Panoramic Depth Anything with Unlab…

200 papers

This paper presents an approach for applying camera perception techniques to spinning LiDAR data. To improve the robustness of long-term change detection from a 3D LiDAR, range and intensity information are rendered into virtual…

Robotics · Computer Science 2024-05-01 Alexander Krawciw , Sven Lilge , Timothy D. Barfoot

The diversity of retinal imaging devices poses a significant challenge: domain shift, which leads to performance degradation when applying the deep learning models trained on one domain to new testing domains. In this paper, we propose a…

Image and Video Processing · Electrical Eng. & Systems 2021-10-07 Peng Liu , Charlie T. Tran , Bin Kong , Ruogu Fang

We introduce a convolutional neural network model for unsupervised learning of depth and ego-motion from cylindrical panoramic video. Panoramic depth estimation is an important technology for applications such as virtual reality, 3D…

Computer Vision and Pattern Recognition · Computer Science 2020-11-11 Alisha Sharma , Ryan Nett , Jonathan Ventura

LiDAR sensors are often considered essential for autonomous driving, but high-resolution sensors remain expensive while affordable low-resolution sensors produce sparse point clouds that miss critical details. LiDAR super-resolution…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 June Moh Goo , Zichao Zeng , Jan Boehm

The Segment Anything Model (SAM) made an eye-catching debut recently and inspired many researchers to explore its potential and limitation in terms of zero-shot generalization capability. As the first promptable foundation model for…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Dongjie Cheng , Ziyuan Qin , Zekun Jiang , Shaoting Zhang , Qicheng Lao , Kang Li

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Haotong Lin , Sili Chen , Junhao Liew , Donny Y. Chen , Zhenyu Li , Guang Shi , Jiashi Feng , Bingyi Kang

Deep learning approaches for semantic segmentation rely primarily on supervised learning approaches and require substantial efforts in producing pixel-level annotations. Further, such approaches may perform poorly when applied to unseen…

Computer Vision and Pattern Recognition · Computer Science 2021-10-22 Ying Chen , Xu Ouyang , Kaiyue Zhu , Gady Agam

Panoramic X-ray is a simple and effective tool for diagnosing dental diseases in clinical practice. When deep learning models are developed to assist dentist in interpreting panoramic X-rays, most of their performance suffers from the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Zijian Cai , Xinquan Yang , Xuguang Li , Xiaoling Luo , Xuechen Li , Linlin Shen , He Meng , Yongqiang Deng

Weakly supervised semantic segmentation (WSSS) aims to bypass the need for laborious pixel-level annotation by using only image-level annotation. Most existing methods rely on Class Activation Maps (CAM) to derive pixel-level pseudo-labels…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Tianle Chen , Zheda Mai , Ruiwen Li , Wei-lun Chao

3D point cloud semantic segmentation (PCSS) is a cornerstone for environmental perception in robotic systems and autonomous driving, enabling precise scene understanding through point-wise classification. While unsupervised domain…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Junjie Chen , Yuecong Xu , Haosheng Li , Kemi Ding

Monocular depth estimation is crucial for tracking and reconstruction algorithms, particularly in the context of surgical videos. However, the inherent challenges in directly obtaining ground truth depth maps during surgery render…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Ange Lou , Yamin Li , Yike Zhang , Jack Noble

Generating detailed and accurate descriptions for specific regions in images and videos remains a fundamental challenge for vision-language models. We introduce the Describe Anything Model (DAM), a model designed for detailed localized…

Computer Vision and Pattern Recognition · Computer Science 2025-04-23 Long Lian , Yifan Ding , Yunhao Ge , Sifei Liu , Hanzi Mao , Boyi Li , Marco Pavone , Ming-Yu Liu , Trevor Darrell , Adam Yala , Yin Cui

Learning semantic representations from point sets of 3D object shapes is often challenged by significant geometric variations, primarily due to differences in data acquisition methods. Typically, training data is generated using point…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Longkun Zou , Kangjun Liu , Ke Chen , Kailing Guo , Kui Jia , Yaowei Wang

Indoor localization in GPS-denied environments is crucial for applications like emergency response and assistive navigation. Vision-based methods such as PALMS enable infrastructure-free localization using only a floor plan and a stationary…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Yunqian Cheng , Benjamin Princen , Roberto Manduchi

In this paper, we present a comprehensive investigation of the challenges of Monocular Visual Simultaneous Localization and Mapping (vSLAM) methods for underwater robots. While significant progress has been made in state estimation methods…

Robotics · Computer Science 2025-07-29 Michele Grimaldi , David Nakath , Mengkun She , Kevin Köser

Self-supervised depth estimation has made a great success in learning depth from unlabeled image sequences. While the mappings between image and pixel-wise depth are well-studied in current methods, the correlation between image, depth and…

Computer Vision and Pattern Recognition · Computer Science 2021-02-15 Rui Li , Xiantuo He , Danna Xue , Shaolin Su , Qing Mao , Yu Zhu , Jinqiu Sun , Yanning Zhang

Model generalizability to unseen datasets, concerned with in-the-wild robustness, is less studied for indoor single-image depth prediction. We leverage gradient-based meta-learning for higher generalizability on zero-shot cross-dataset…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Cho-Ying Wu , Yiqi Zhong , Junying Wang , Ulrich Neumann

Computing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision. However, most existing approaches do not generalize well across different domains and scene types (e.g. indoor vs. outdoor).…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Sergio Izquierdo , Mohamed Sayed , Michael Firman , Guillermo Garcia-Hernando , Daniyar Turmukhambetov , Javier Civera , Oisin Mac Aodha , Gabriel Brostow , Jamie Watson

Recently, few-shot learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Liwen Wu , Wei Wang , Lei Zhao , Zhan Gao , Qika Lin , Shaowen Yao , Zuozhu Liu , Bin Pu

The Segment Anything Model (SAM) marks a significant advancement in segmentation models, offering robust zero-shot abilities and dynamic prompting. However, existing medical SAMs are not suitable for the multi-scale nature of whole-slide…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Hong Liu , Haosen Yang , Paul J. van Diest , Josien P. W. Pluim , Mitko Veta