中文
相关论文

相关论文: V$^2$-SfMLearner: Learning Monocular Depth and Ego…

200 篇论文

Accurate endoscope pose estimation and 3D tissue surface reconstruction significantly enhances monocular minimally invasive surgical procedures by enabling accurate navigation and improved spatial awareness. However, monocular endoscope…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Muzammil Khan , Enzo Kerkhof , Matteo Fusaglia , Koert Kuhlmann , Theo Ruers , Françoise J. Siepel

Artificial intelligence holds strong potential to support clinical decision making in intensive care units where timely and accurate risk assessment is critical. However, many existing models focus on isolated outcomes or limited data…

Due to its non-invasive and painless characteristics, wireless capsule endoscopy has become the new gold standard for assessing gastrointestinal disorders. Omissions, however, could occur throughout the examination since controlling capsule…

机器人学 · 计算机科学 2023-05-19 Yameng Zhang , Long Bai , Li Liu , Hongliang Ren , Max Q. -H. Meng

Reconstructing 3D scenes from monocular surgical videos can enhance surgeon's perception and therefore plays a vital role in various computer-assisted surgery tasks. However, achieving scale-consistent reconstruction remains an open…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Jiaxin Guo , Wenzhen Dong , Tianyu Huang , Hao Ding , Ziyi Wang , Haomin Kuang , Qi Dou , Yun-Hui Liu

Purpose: Monocular depth estimation (MDE) is vital for scene understanding in minimally invasive surgery (MIS). However, endoscopic video sequences are often contaminated by smoke, specular reflections, blur, and occlusions, limiting the…

Gastroendoscopy has been a clinical standard for diagnosing and treating conditions that affect a part of a patient's digestive system, such as the stomach. Despite the fact that gastroendoscopy has a lot of advantages for patients, there…

计算机视觉与模式识别 · 计算机科学 2021-07-29 Aji Resindra Widya , Yusuke Monno , Masatoshi Okutomi , Sho Suzuki , Takuji Gotoda , Kenji Miki

Dense depth estimation is essential to scene-understanding for autonomous driving. However, recent self-supervised approaches on monocular videos suffer from scale-inconsistency across long sequences. Utilizing data from the ubiquitously…

计算机视觉与模式识别 · 计算机科学 2023-02-03 Hemang Chawla , Arnav Varma , Elahe Arani , Bahram Zonooz

Recently, deep learning-based tooth segmentation methods have been limited by the expensive and time-consuming processes of data collection and labeling. Achieving high-precision segmentation with limited datasets is critical. A viable…

计算机视觉与模式识别 · 计算机科学 2023-10-24 Yuan Li , Huan Liu , Yubo Tao , Xiangyang He , Haifeng Li , Xiaohu Guo , Hai Lin

Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often struggle in scenarios requiring precise spatial…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Tao Lin , Yuxin Du , Jiting Liu , Nuobei Zhu , Yunhe Li , Yuqian Fu , Yinxinyu Chen , Hongyi Cai , Zewei Ye , Bing Cheng , Kai Ye , Yiran Mao , Yilei Zhong , MingKang Dong , Junchi Yan , Gen Li , Bo Zhao

The field of indoor monocular 3D object detection is gaining significant attention, fueled by the increasing demand in VR/AR and robotic applications. However, its advancement is impeded by the limited availability and diversity of 3D…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Jin-Cheng Jhang , Tao Tu , Fu-En Wang , Ke Zhang , Min Sun , Cheng-Hao Kuo

The scalability of robotic manipulation is fundamentally bottlenecked by the scarcity of task-aligned physical interaction data. While vision-language models (VLMs) and video generation models (VGMs) hold promise for autonomous data…

机器人学 · 计算机科学 2026-05-14 Harold Haodong Chen , Sirui Chen , Yingjie Xu , Wenhang Ge , Ying-Cong Chen

We present an approach which takes advantage of both structure and semantics for unsupervised monocular learning of depth and ego-motion. More specifically, we model the motion of individual objects and learn their 3D motion vector jointly…

计算机视觉与模式识别 · 计算机科学 2019-06-14 Vincent Casser , Soeren Pirk , Reza Mahjourian , Anelia Angelova

Purpose: Surgical scene understanding plays a critical role in the technology stack of tomorrow's intervention-assisting systems in endoscopic surgeries. For this, tracking the endoscope pose is a key component, but remains challenging due…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Michel Hayoz , Christopher Hahne , Mathias Gallardo , Daniel Candinas , Thomas Kurmann , Maximilian Allan , Raphael Sznitman

Egocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key challenges: 1) Excessive pre-training cost arising from multi-stage pre-training pipelines, 2)…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Xiaoqi Wang , Yi Wang , Lap-Pui Chau

Wireless Capsule Endoscopy (WCE) helps physicians examine the gastrointestinal (GI) tract noninvasively. There are few studies that address pathological assessment of endoscopy images in multiclass classification and most of them are based…

计算机视觉与模式识别 · 计算机科学 2021-08-23 Mohammad Reza Mohebbian , Khan A. Wahid , Paul Babyn

BACKGROUND AND PURPOSE: Deep learning has been demonstrated effective in many neuroimaging applications. However, in many scenarios, the number of imaging sequences capturing information related to small vessel disease lesions is…

Endoscopy is a routine imaging technique used for both diagnosis and minimally invasive surgical treatment. While the endoscopy video contains a wealth of information, tools to capture this information for the purpose of clinical reporting…

计算机视觉与模式识别 · 计算机科学 2019-05-14 Sharib Ali , Jens Rittscher

Depth estimation is a cornerstone of 3D reconstruction and plays a vital role in minimally invasive endoscopic surgeries. However, most current depth estimation networks rely on traditional convolutional neural networks, which are limited…

计算机视觉与模式识别 · 计算机科学 2025-07-16 Bojian Li , Bo Liu , Xinning Yao , Jinghua Yue , Fugen Zhou

Medical endoscopy remains a challenging application for simultaneous localization and mapping (SLAM) due to the sparsity of image features and size constraints that prevent direct depth-sensing. We present a SLAM approach that incorporates…

图像与视频处理 · 电气工程与系统科学 2019-07-02 Richard J. Chen , Taylor L. Bobrow , Thomas Athey , Faisal Mahmood , Nicholas J. Durr

Effective and rapid detection of lesions in the Gastrointestinal tract is critical to gastroenterologist's response to some life-threatening diseases. Wireless Capsule Endoscopy (WCE) has revolutionized traditional endoscopy procedure by…

计算机视觉与模式识别 · 计算机科学 2021-01-19 Sodiq Adewole , Philip Fernandez , Michelle Yeghyayan , James Jablonski , Andrew Copland , Michael Porter , Sana Syed , Donald Brown