English
Related papers

Related papers: Extraction of Key-frames of Endoscopic Videos by u…

200 papers

Monocular Depth Estimation (MDE) enables spatial understanding, 3D reconstruction, and autonomous navigation, yet deep learning approaches often predict only relative depth without a consistent metric scale. This limitation reduces…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Jiuling Zhang

Previous methods on estimating detailed human depth often require supervised training with `ground truth' depth data. This paper presents a self-supervised method that can be trained on YouTube videos without known depth, which makes…

Computer Vision and Pattern Recognition · Computer Science 2020-05-08 Feitong Tan , Hao Zhu , Zhaopeng Cui , Siyu Zhu , Marc Pollefeys , Ping Tan

Monocular depth reconstruction of complex and dynamic scenes is a highly challenging problem. While for rigid scenes learning-based methods have been offering promising results even in unsupervised cases, there exists little to no…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Ayça Takmaz , Danda Pani Paudel , Thomas Probst , Ajad Chhatkuli , Martin R. Oswald , Luc Van Gool

Objective: To develop and validate a deep learning model for the identification of out-of-body images in endoscopic videos. Background: Surgical video analysis facilitates education and research. However, video recordings of endoscopic…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Joël L. Lavanchy , Armine Vardazaryan , Pietro Mascagni , AI4SafeChole Consortium , Didier Mutter , Nicolas Padoy

Compression is essential to storing and transmitting medical videos, but the effect of compression on downstream medical tasks is often ignored. Furthermore, systems in practice rely on standard video codecs, which naively allocate bits…

Image and Video Processing · Electrical Eng. & Systems 2022-11-04 Joel Shor , Nick Johnston

Current self-supervised monocular depth estimation methods are mostly based on estimating a rigid-body motion representing camera motion. These methods suffer from the well-known scale ambiguity problem in their predictions. We propose…

Computer Vision and Pattern Recognition · Computer Science 2023-01-06 Sadra Safadoust , Fatma Güney

Estimating depth from single RGB images and videos is of widespread interest due to its applications in many areas, including autonomous driving, 3D reconstruction, digital entertainment, and robotics. More than 500 deep learning-based…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Uchitha Rajapaksha , Ferdous Sohel , Hamid Laga , Dean Diepeveen , Mohammed Bennamoun

This paper presents a probabilistic approach for online dense reconstruction using a single monocular camera moving through the environment. Compared to spatial stereo, depth estimation from motion stereo is challenging due to insufficient…

Robotics · Computer Science 2019-03-27 Yonggen Ling , Kaixuan Wang , Shaojie Shen

One of the major challenges in Minimally Invasive Surgery (MIS) such as laparoscopy is the lack of depth perception. In recent years, laparoscopic scene tracking and surface reconstruction has been a focus of investigation to provide rich…

Computer Vision and Pattern Recognition · Computer Science 2017-03-06 Long Chen , Wen Tang , Nigel W. John , Tao Ruan Wan , Jian Jun Zhang

Localizing oneself during endoscopic procedures can be problematic due to the lack of distinguishable textures and landmarks, as well as difficulties due to the endoscopic device such as a limited field of view and challenging lighting…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Gary Sarwin , Alessandro Carretta , Victor Staartjes , Matteo Zoli , Diego Mazzatenta , Luca Regli , Carlo Serra , Ender Konukoglu

Self-supervised monocular methods can efficiently learn depth information of weakly textured surfaces or reflective objects. However, the depth accuracy is limited due to the inherent ambiguity in monocular geometric modeling. In contrast,…

Computer Vision and Pattern Recognition · Computer Science 2022-08-22 Xiaofeng Wang , Zheng Zhu , Guan Huang , Xu Chi , Yun Ye , Ziwei Chen , Xingang Wang

Multimodal large language models (MLLMs) represent images and video frames as visual tokens. Scaling from single images to hour-long videos, however, inflates the token budget far beyond practical limits. Popular pipelines therefore either…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Zirui Zhu , Hailun Xu , Yang Luo , Yong Liu , Kanchan Sarkar , Zhenheng Yang , Yang You

For ego-motion estimation, the feature representation of the scenes is crucial. Previous methods indicate that both the low-level and semantic feature-based methods can achieve promising results. Therefore, the incorporation of hierarchical…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Xiaochuan Yin , Chengju Liu

Endoscopic video recordings are widely used in minimally invasive robot-assisted surgery, but when the endoscope is outside the patient's body, it can capture irrelevant segments that may contain sensitive information. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2023-04-03 Ziheng Wang , Conor Perreault , Xi Liu , Anthony Jarc

Perceiving 3D information is of paramount importance in many applications of computer vision. Recent advances in monocular depth estimation have shown that gaining such knowledge from a single camera input is possible by training deep…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Sai Shyam Chanduri , Zeeshan Khan Suri , Igor Vozniak , Christian Müller

We introduce the task of stereo video reconstruction or, equivalently, 2D-to-3D video conversion for minimally invasive surgical video. We design and implement a series of end-to-end U-Net-based solutions for this task by varying the input…

Image and Video Processing · Electrical Eng. & Systems 2021-09-20 Annika Brundyn , Jesse Swanson , Kyunghyun Cho , Doug Kondziolka , Eric Oermann

Deep learning techniques hold promise to develop dense topography reconstruction and pose estimation methods for endoscopic videos. However, currently available datasets do not support effective quantitative benchmarking. In this paper, we…

While a traditional camera only captures one point of view of a scene, a plenoptic or light-field camera, is able to capture spatial and angular information in a single snapshot, enabling depth estimation from a single acquisition. In this…

Image and Video Processing · Electrical Eng. & Systems 2023-08-09 Mathieu Labussière , Céline Teulière , Omar Ait-Aider

Pre-training on image-text colonoscopy records offers substantial potential for improving endoscopic image analysis, but faces challenges including non-informative background images, complex medical terminology, and ambiguous multi-lesion…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Yili He , Yan Zhu , Peiyao Fu , Ruijie Yang , Tianyi Chen , Zhihua Wang , Quanlin Li , Pinghong Zhou , Xian Yang , Shuo Wang

Monocular Depth Estimation (MDE) is a fundamental computer vision task with important applications in 3D vision. The current mainstream MDE methods employ an encoder-decoder architecture with multi-level/scale feature processing. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Huibin Bai , Shuai Li , Hanxiao Zhai , Yanbo Gao , Chong Lv , Yibo Wang , Haipeng Ping , Wei Hua , Xingyu Gao
‹ Prev 1 3 4 5 6 7 10 Next ›