English
Related papers

Related papers: EndoDINO: A Foundation Model for GI Endoscopy

200 papers

We present DINO Patch Visual Odometry (DINO-VO), an end-to-end monocular visual odometry system with strong scene generalization. Current Visual Odometry (VO) systems often rely on heuristic feature extraction strategies, which can degrade…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Qi Chen , Guanghao Li , Sijia Hu , Xin Gao , Junpeng Ma , Xiangyang Xue , Jian Pu

Reliable and real-time 3D reconstruction and localization functionality is a crucial prerequisite for the navigation of actively controlled capsule endoscopic robots as an emerging, minimally invasive diagnostic and therapeutic technology…

Recent advancements in foundation models have transformed computer vision, driving significant performance improvements across diverse domains, including digital histopathology. However, the advantages of domain-specific histopathology…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Valentina Vadori , Antonella Peruffo , Jean-Marie Graïc , Livio Finos , Enrico Grisan

Gastrointestinal (GI) diseases represent a clinically significant burden, necessitating precise diagnostic approaches to optimize patient outcomes. Conventional histopathological diagnosis suffers from limited reproducibility and diagnostic…

Endoscopic procedures are essential for diagnosing and treating internal diseases, and multi-modal large language models (MLLMs) are increasingly applied to assist in endoscopy analysis. However, current benchmarks are limited, as they…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Shengyuan Liu , Boyun Zheng , Wenting Chen , Zhihao Peng , Zhenfei Yin , Jing Shao , Jiancong Hu , Yixuan Yuan

There has been a longstanding expectation that the optical resolution embodiment of photoacoustic tomography could have a substantial impact on gastrointestinal endoscopy by enabling microscopic visualization of the vasculature based on the…

Vision Transformers (ViTs) have demonstrated strong potential in medical imaging; however, their high computational demands and tendency to overfit on small datasets limit their applicability in real-world clinical scenarios. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Aon Safdar , Mohamed Saadeldin

Existing foundation models (FMs) in the medical domain often require extensive fine-tuning or rely on training resource-intensive decoders, while many existing encoders are pretrained with objectives biased toward specific tasks. This…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Tim Veenboer , George Yiasemis , Eric Marcus , Vivien Van Veldhuizen , Cees G. M. Snoek , Jonas Teuwen , Kevin B. W. Groot Lipman

Automatic polyp segmentation is crucial for improving the clinical identification of colorectal cancer (CRC). While Deep Learning (DL) techniques have been extensively researched for this problem, current methods frequently struggle with…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Carla Monteiro , Valentina Corbetta , Regina Beets-Tan , Luís F. Teixeira , Wilson Silva

Colonoscopy is a routine outpatient procedure used to examine the colon and rectum for any abnormalities including polyps, diverticula and narrowing of colon structures. A significant amount of the clinician's time is spent in…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Aniruddha Tamhane , Tse'ela Mida , Erez Posner , Moshe Bouhnik

Accurate and robust localization is a fundamental need for mobile agents. Visual-inertial odometry (VIO) algorithms exploit the information from camera and inertial sensors to estimate position and translation. Recent deep learning based…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Zheming Tu , Changhao Chen , Xianfei Pan , Ruochen Liu , Jiarui Cui , Jun Mao

Continuous monitoring of foot ulcer healing is needed to ensure the efficacy of a given treatment and to avoid any possibility of deterioration. Foot ulcer segmentation is an essential step in wound diagnosis. We developed a model that is…

Image and Video Processing · Electrical Eng. & Systems 2022-07-07 Shahzad Ali , Arif Mahmood , Soon Ki Jung

Image-based tracking of medical instruments is an integral part of surgical data science applications. Previous research has addressed the tasks of detecting, segmenting and tracking medical instruments based on laparoscopic video data.…

Purpose: Surgical scene understanding plays a critical role in the technology stack of tomorrow's intervention-assisting systems in endoscopic surgeries. For this, tracking the endoscope pose is a key component, but remains challenging due…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Michel Hayoz , Christopher Hahne , Mathias Gallardo , Daniel Candinas , Thomas Kurmann , Maximilian Allan , Raphael Sznitman

Deformable image registration and regression are important tasks in medical image analysis. However, they are computationally expensive, especially when analyzing large-scale datasets that contain thousands of images. Hence, cluster…

Computer Vision and Pattern Recognition · Computer Science 2017-11-17 Zhipeng Ding , Greg Fleishman , Xiao Yang , Paul Thompson , Roland Kwitt , Marc Niethammer

Automating the analysis of imagery of the Gastrointestinal (GI) tract captured during endoscopy procedures has substantial potential benefits for patients, as it can provide diagnostic support to medical practitioners and reduce mistakes…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Harshala Gammulle , Simon Denman , Sridha Sridharan , Clinton Fookes

Self-supervised monocular depth estimation is a significant task for low-cost and efficient 3D scene perception and measurement in endoscopy. However, the variety of illumination conditions and scene features is still the primary challenges…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Liangjing Shao , Chenkang Du , Benshuang Chen , Xueli Liu , Xinrong Chen

This paper proposes novel methods to enhance the performance of monocular 3D object detection models by leveraging the generalized feature extraction capabilities of a vision foundation model. Unlike traditional CNN-based approaches, which…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Jihyeok Kim , Seongwoo Moon , Sungwon Nah , David Hyunchul Shim

Gastrointestinal (GI) diseases represent a significant global health concern, with Capsule Endoscopy (CE) offering a non-invasive method for diagnosis by capturing a large number of GI tract images. However, the sheer volume of video frames…

Image and Video Processing · Electrical Eng. & Systems 2024-10-28 Aniket Das , Ayushman Singh , Nishant , Sharad Prakash

Compared to the great progress of large-scale vision transformers (ViTs) in recent years, large-scale models based on convolutional neural networks (CNNs) are still in an early state. This work presents a new large-scale CNN-based…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Wenhai Wang , Jifeng Dai , Zhe Chen , Zhenhang Huang , Zhiqi Li , Xizhou Zhu , Xiaowei Hu , Tong Lu , Lewei Lu , Hongsheng Li , Xiaogang Wang , Yu Qiao