中文
相关论文

相关论文: Harnessing Foundation Models for Robust and Genera…

200 篇论文

Accurate bronchoscope localization is essential for pulmonary interventions, by providing six degrees of freedom (DOF) in airway navigation. However, the robustness of current vision-based methods is often compromised in clinical practice,…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Qingyao Tian , Zhen Chen , Huai Liao , Xinyan Huang , Bingyu Yang , Lujie Li , Hongbin Liu

Real-time 6 DOF localization of bronchoscopes is crucial for enhancing intervention quality. However, current vision-based technologies struggle to balance between generalization to unseen data and computational speed. In this study, we…

计算机视觉与模式识别 · 计算机科学 2025-02-12 Qingyao Tian , Huai Liao , Xinyan Huang , Jian Chen , Zihui Zhang , Bingyu Yang , Sebastien Ourselin , Hongbin Liu

Accurate intra-operative localization of the bronchoscope tip relative to patient anatomy remains challenging due to respiratory motion, anatomical variability, and CT-to-body divergence that cause deformation and misalignment between…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Hongchao Shu , Roger D. Soberanis-Mukul , Jiru Xu , Hao Ding , Morgan Ringel , Mali Shen , Saif Iftekar Sayed , Hedyeh Rafii-Tari , Mathias Unberath

Learning-based monocular visual odometry (VO) poses robustness, generalization, and efficiency challenges in robotics. Recent advances in visual foundation models, such as DINOv2, have improved robustness and generalization in various…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Maulana Bisyir Azhari , David Hyunchul Shim

Vision-language models (VLMs) have recently shown remarkable performance in navigation and localization tasks by leveraging large-scale pretraining for semantic understanding. However, applying VLMs to 6-DoF endoscopic camera localization…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Qingyao Tian , Bingyu Yang , Huai Liao , Xinyan Huang , Junyong Li , Dong Yi , Hongbin Liu

3D reconstruction of endoscopic surgery scenes plays a vital role in enhancing scene perception, enabling AR visualization, and supporting context-aware decision-making in image-guided surgery. A critical yet challenging step in this…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Changhao Zhang , Matthew J. Clarkson , Mobarak I. Hoque

Depth estimation is a foundational component for 3D reconstruction in minimally invasive endoscopic surgeries. However, existing monocular depth estimation techniques often exhibit limited performance to the varying illumination and complex…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Xinning Yao , Bo Liu , Bojian Li , Jingjing Wang , Jinghua Yue , Fugen Zhou

Localization in a battlefield environment is increasingly challenging as GPS connectivity is often denied or unreliable, and physical deployment of anchor nodes across wireless networks for localization can be difficult in hostile…

计算机视觉与模式识别 · 计算机科学 2024-02-21 Ganesh Sapkota , Sanjay Madria

The integration of deep learning systems into healthcare has been hindered by the resource-intensive process of data annotation and the inability of these systems to generalize to different data distributions. Foundation models, which are…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Mohammed Baharoon , Waseem Qureshi , Jiahong Ouyang , Yanwu Xu , Abdulrhman Aljouie , Wei Peng

Bird's-Eye-View (BEV) representation offers a metric-scaled planar workspace, facilitating the simplification of 6-DoF ego-motion to a more robust 3-DoF model for monocular visual odometry (MVO) in intelligent transportation systems.…

机器人学 · 计算机科学 2025-09-19 Yufei Wei , Wangtao Lu , Sha Lu , Chenxiao Hu , Fuzhang Han , Rong Xiong , Yue Wang

Precise six-degree-of-freedom (6DoF) head pose estimation is crucial for safety-critical applications and human-computer interaction scenarios, yet existing monocular methods still struggle with robust pose estimation. We revisit this…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Sungho Chun , Boeun Kim , Hyung Jin Chang , Ju Yong Chang

Fetal pose estimation in 3D ultrasound (US) involves identifying a set of associated fetal anatomical landmarks. Its primary objective is to provide comprehensive information about the fetus through landmark connections, thus benefiting…

图像与视频处理 · 电气工程与系统科学 2023-10-31 Chaoyu Chen , Xin Yang , Yuhao Huang , Wenlong Shi , Yan Cao , Mingyuan Luo , Xindi Hu , Lei Zhue , Lequan Yu , Kejuan Yue , Yuanji Zhang , Yi Xiong , Dong Ni , Weijun Huang

Vision foundation models like DINOv2 demonstrate remarkable potential in medical imaging despite their origin in natural image domains. However, their design inherently works best for uni-modal image analysis, limiting their effectiveness…

图像与视频处理 · 电气工程与系统科学 2025-09-09 Daniel Scholz , Ayhan Can Erdur , Viktoria Ehm , Anke Meyer-Baese , Jan C. Peeken , Daniel Rueckert , Benedikt Wiestler

Purpose: Surgical scene understanding plays a critical role in the technology stack of tomorrow's intervention-assisting systems in endoscopic surgeries. For this, tracking the endoscope pose is a key component, but remains challenging due…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Michel Hayoz , Christopher Hahne , Mathias Gallardo , Daniel Candinas , Thomas Kurmann , Maximilian Allan , Raphael Sznitman

Localizing the bronchoscope in real time is essential for ensuring intervention quality. However, most existing methods struggle to balance between speed and generalization. To address these challenges, we present BronchoTrack, an…

计算机视觉与模式识别 · 计算机科学 2024-02-21 Qingyao Tian , Huai Liao , Xinyan Huang , Bingyu Yang , Jinlin Wu , Jian Chen , Lujie Li , Hongbin Liu

Depth estimation plays a crucial role in various tasks within endoscopic surgery, including navigation, surface reconstruction, and augmented reality visualization. Despite the significant achievements of foundation models in vision tasks,…

图像与视频处理 · 电气工程与系统科学 2024-05-15 Beilei Cui , Mobarakol Islam , Long Bai , An Wang , Hongliang Ren

We present FoundationPose, a unified foundation model for 6D object pose estimation and tracking, supporting both model-based and model-free setups. Our approach can be instantly applied at test-time to a novel object without fine-tuning,…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Bowen Wen , Wei Yang , Jan Kautz , Stan Birchfield

Vision-based bronchoscopy (VB) models require the registration of the virtual lung model with the frames from the video bronchoscopy to provide effective guidance during the biopsy. The registration can be achieved by either tracking the…

计算机视觉与模式识别 · 计算机科学 2022-04-27 Juan Borrego-Carazo , Carles Sánchez , David Castells-Rufas , Jordi Carrabina , Débora Gil

Visual localization has become a key enabling component of many place recognition and SLAM systems. Contemporary research has primarily focused on improving accuracy and precision-recall type metrics, with relatively little attention paid…

计算机视觉与模式识别 · 计算机科学 2019-02-06 Huu Le , Tuan Hoang , Qianggong Zhang , Thanh-Toan Do , Anders Eriksson , Michael Milford

Gastroendoscopy has been a clinical standard for diagnosing and treating conditions that affect a part of a patient's digestive system, such as the stomach. Despite the fact that gastroendoscopy has a lot of advantages for patients, there…

计算机视觉与模式识别 · 计算机科学 2021-07-29 Aji Resindra Widya , Yusuke Monno , Masatoshi Okutomi , Sho Suzuki , Takuji Gotoda , Kenji Miki
‹ 上一页 1 2 3 10 下一页 ›