English
Related papers

Related papers: SLAM Endoscopy enhanced by adversarial depth predi…

200 papers

This work presents EndoStreamDepth, a monocular depth estimation framework for endoscopic video streams. It provides accurate depth maps with sharp anatomical boundaries for each frame, temporally consistent predictions across frames, and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Hao Li , Daiwei Lu , Jiacheng Wang , Robert J. Webster , Ipek Oguz

In this work, we develop a monocular SLAM-aware object recognition system that is able to achieve considerably stronger recognition performance, as compared to classical object recognition systems that function on a frame-by-frame basis. By…

Robotics · Computer Science 2015-06-08 Sudeep Pillai , John Leonard

To enhance the performance and effect of AR/VR applications and visual assistance and inspection systems, visual simultaneous localization and mapping (vSLAM) is a fundamental task in computer vision and robotics. However, traditional vSLAM…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 Yichen Chen , Yiqi Pan , Ruyu Liu , Haoyu Zhang , Guodao Zhang , Bo Sun , Jianhua Zhang

Gaussian splatting has recently gained traction as a compelling map representation for SLAM systems, enabling dense and photo-realistic scene modeling. However, its application to monocular SLAM remains challenging due to the lack of…

Robotics · Computer Science 2026-04-20 Dong-Uk Seo , Jinwoo Jeon , Eungchang Mason Lee , Hyun Myung

We propose a depth estimation method from a single-shot monocular endoscopic image using Lambertian surface translation by domain adaptation and depth estimation using multi-scale edge loss. We employ a two-step estimation process including…

Image and Video Processing · Electrical Eng. & Systems 2022-01-13 Masahiro Oda , Hayato Itoh , Kiyohito Tanaka , Hirotsugu Takabatake , Masaki Mori , Hiroshi Natori , Kensaku Mori

Monocular depth estimation (MDE) for colonoscopy is hampered by the domain gap between simulated and real-world images. Existing image-to-image translation methods, which use depth as a posterior constraint, often produce structural…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Juan Yang , Yuyan Zhang , Han Jia , Bing Hu , Wanzhong Song

This paper proposes an adversarial attack method to deep neural networks (DNNs) for monocular depth estimation, i.e., estimating the depth from a single image. Single image depth estimation has improved drastically in recent years due to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Renya Daimo , Satoshi Ono , Takahiro Suzuki

High screening coverage during colonoscopy is crucial to effectively prevent colon cancer. Previous work has allowed alerting the doctor to unsurveyed regions by reconstructing the 3D colonoscopic surface from colonoscopy videos in…

Computer Vision and Pattern Recognition · Computer Science 2021-03-19 Yubo Zhang , Shuxian Wang , Ruibin Ma , Sarah K. McGill , Julian G. Rosenman , Stephen M. Pizer

Recently, convolutional neural networks (CNNs) have shown great success on the task of monocular depth estimation. A fundamental yet unanswered question is: how CNNs can infer depth from a single image. Toward answering this question, we…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Junjie Hu , Yan Zhang , Takayuki Okatani

Neural implicit representations have recently become popular in simultaneous localization and mapping (SLAM), especially in dense visual SLAM. However, previous works in this direction either rely on RGB-D sensors, or require a separate…

Computer Vision and Pattern Recognition · Computer Science 2023-02-08 Zihan Zhu , Songyou Peng , Viktor Larsson , Zhaopeng Cui , Martin R. Oswald , Andreas Geiger , Marc Pollefeys

The advent of deep learning has brought an impressive advance to monocular depth estimation, e.g., supervised monocular depth estimation has been thoroughly investigated. However, the large amount of the RGB-to-depth dataset may not be…

Computer Vision and Pattern Recognition · Computer Science 2021-04-14 Fei Lu , Hyeonwoo Yu , Jean Oh

We propose a novel dense mapping framework for sparse visual SLAM systems which leverages a compact scene representation. State-of-the-art sparse visual SLAM systems provide accurate and reliable estimates of the camera trajectory and…

Computer Vision and Pattern Recognition · Computer Science 2021-07-20 Hidenobu Matsuki , Raluca Scona , Jan Czarnowski , Andrew J. Davison

In order to use the navigation system effectively, distance information sensors such as depth sensors are essential. Since depth sensors are difficult to use in endoscopy, many groups propose a method using convolutional neural networks. In…

Image and Video Processing · Electrical Eng. & Systems 2021-12-28 Bong Hyuk Jeong , Hang Keun Kim , Young Don Son

Autonomous navigation in GPS-denied and visually degraded environments remains challenging for unmanned aerial vehicles (UAVs). To this end, we investigate the use of a monocular thermal camera as a standalone sensor on a UAV platform for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Hürkan Şahin , Huy Xuan Pham , Van Huyen Dang , Alper Yegenoglu , Erdal Kayacan

Monocular SLAM refers to using a single camera to estimate robot ego motion while building a map of the environment. While Monocular SLAM is a well studied problem, automating Monocular SLAM by integrating it with trajectory planning…

Monocular SLAM refers to using a single camera to estimate robot ego motion while building a map of the environment. While Monocular SLAM is a well studied problem, automating Monocular SLAM by integrating it with trajectory planning…

We present the first application of 3D Gaussian Splatting in monocular SLAM, the most fundamental but the hardest setup for Visual SLAM. Our method, which runs live at 3fps, utilises Gaussians as the only 3D representation, unifying the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Hidenobu Matsuki , Riku Murai , Paul H. J. Kelly , Andrew J. Davison

Medical Visual Question Answering (Medical-VQA) aims to to answer clinical questions regarding radiology images, assisting doctors with decision-making options. Nevertheless, current Medical-VQA models learn cross-modal representations…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Chenlu Zhan , Peng Peng , Hongsen Wang , Tao Chen , Hongwei Wang

Estimating depth information from endoscopic images is a prerequisite for a wide set of AI-assisted technologies, such as accurate localization and measurement of tumors, or identification of non-inspected areas. As the domain specificity…

Computer Vision and Pattern Recognition · Computer Science 2022-07-21 Javier Rodríguez-Puigvert , David Recasens , Javier Civera , Rubén Martínez-Cantín

This work delves into unsupervised monocular depth estimation in endoscopy, which leverages adjacent frames to establish a supervisory signal during the training phase. For many clinical applications, e.g., surgical navigation, temporally…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Shuwei Shao , Zhongcai Pei , Weihai Chen , Xingming Wu , Zhong Liu