English
Related papers

Related papers: Endo-FASt3r: Endoscopic Foundation model Adaptatio…

200 papers

Estimating relative camera poses between images has been a central problem in computer vision. Methods that find correspondences and solve for the fundamental matrix offer high precision in most cases. Conversely, methods predicting pose…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Chris Rockwell , Nilesh Kulkarni , Linyi Jin , Jeong Joon Park , Justin Johnson , David F. Fouhey

Continuum manipulators in flexible endoscopic surgical systems offer high dexterity for minimally invasive procedures; however, accurate pose estimation and closed-loop control remain challenging due to hysteresis, compliance, and limited…

Robotics · Computer Science 2026-02-19 Junhyun Park , Chunggil An , Myeongbo Park , Ihsan Ullah , Sihyeong Park , Minho Hwang

While recent feed-forward 3D reconstruction models provide a strong geometric foundation for scene understanding, extending them to 3D instance segmentation typically relies on a disjointed "lift-and-cluster" paradigm. Grouping dense…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Changyang Li , Xueqing Huang , Shin-Fang Chng , Huangying Zhan , Qingan Yan , Yi Xu

Endoscopic (endo) video exhibits strong view-dependent effects such as specularities, wet reflections, and occlusions. Pure photometric supervision misaligns with geometry and triggers early geometric drift, where erroneous shapes are…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Yangle Liu , Fengze Li , Kan Liu , Jieming Ma

We aim to track the endoscope location inside the surgical scene and provide 3D reconstruction, in real-time, from the sole input of the image sequence captured by the monocular endoscope. This information offers new possibilities for…

Computer Vision and Pattern Recognition · Computer Science 2016-08-30 Nader Mahmoud , Iñigo Cirauqui , Alexandre Hostettler , Christophe Doignon , Luc Soler , Jacques Marescaux , J. M. M. Montiel

The task of 3D semantic scene completion using monocular cameras is gaining significant attention in the field of autonomous driving. This task aims to predict the occupancy status and semantic labels of each voxel in a 3D scene from…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Jiawei Yao , Jusheng Zhang , Xiaochao Pan , Tong Wu , Canran Xiao

In robot-assisted minimally invasive surgery, accurate 3D reconstruction from endoscopic video is vital for downstream tasks and improved outcomes. However, endoscopic scenarios present unique challenges, including photometric…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Taoyu Wu , Yiyi Miao , Jiaxin Guo , Ziyan Chen , Sihang Zhao , Zhuoxiao Li , Zhe Tang , Baoru Huang , Limin Yu

Depth sensing is an important problem for 3D vision-based robotics. Yet, a real-world active stereo or ToF depth camera often produces noisy and incomplete depth which bottlenecks robot performances. In this work, we propose D3RoMa, a…

On-orbit servicing and active debris removal involving non-cooperative spacecraft require reliable pose estimation to supply accurate position and orientation data for autonomous visual navigation. Learning-based monocular methods have seen…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Yongliang Zhen , Bo LÜ , Hang Yang , Xiaotian WU

In this paper, we focus on online zero-shot monocular 3D instance segmentation, a novel practical setting where existing approaches fail to perform because they rely on posed RGB-D sequences. To overcome this limitation, we leverage CUT3R,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Zhipeng Du , Duolikun Danier , Jan Eric Lenssen , Hakan Bilen

Achieving high-fidelity 3D reconstruction from monocular video remains challenging due to the inherent limitations of traditional methods like Structure-from-Motion (SfM) and monocular SLAM in accurately capturing scene details. While…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Yue Hu , Rong Liu , Meida Chen , Peter Beerel , Andrew Feng

Building a robust perception module is crucial for visuomotor policy learning. While recent methods incorporate pre-trained 2D foundation models into robotic perception modules to leverage their strong semantic understanding, they struggle…

Robotics · Computer Science 2025-07-14 Wenbo Cui , Chengyang Zhao , Yuhui Chen , Haoran Li , Zhizheng Zhang , Dongbin Zhao , He Wang

The scarcity and high cost of expert annotations in dental imaging present a significant challenge for the development of AI in dentistry. DINOv3, a state-of-the-art, self-supervised vision foundation model pre-trained on 1.7 billion…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Kun Tang , Xinquan Yang , Mianjie Zheng , Xuefen Liu , Xuguang Li , Xiaoqi Guo , Ruihan Chen , Linlin Shen , He Meng

Panoramic semantic segmentation models are typically trained under a strict gravity-aligned assumption. However, real-world captures often deviate from this canonical orientation due to unconstrained camera motions, such as the rotational…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Qinfeng Zhu , Yunxi Jiang , Lei Fan

Accurate 6-DoF pose estimation of surgical instruments during minimally invasive surgeries can substantially improve treatment strategies and eventual surgical outcome. Existing deep learning methods have achieved accurate results, but they…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Christiaan G. A. Viviers , Lena Filatova , Maurice Termeer , Peter H. N. de With , Fons van der Sommen

Confocal laser endomicroscopy (CLE) is a non-invasive, real-time imaging modality that can be used for in-situ, in-vivo imaging and the microstructural analysis of mucous structures. The diagnosis using CLE is, however, complicated by…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Nils Porsche , Flurin Müller-Diesing , Sweta Banerjee , Miguel Goncalves , Marc Aubreville

Recent 6D pose estimation methods demonstrate notable performance but still face some practical limitations. For instance, many of them rely heavily on sensor depth, which may fail with challenging surface conditions, such as transparent or…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Jiahui Wang , Haiyue Zhu , Haoren Guo , Abdullah Al Mamun , Cheng Xiang , Tong Heng Lee

Existing monocular 3D pose estimation methods primarily rely on joint positional features, while overlooking intrinsic directional and angular correlations within the skeleton. As a result, they often produce implausible poses under joint…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Ming Xu , Xu Zhang

Most 3d human pose estimation methods assume that input -- be it images of a scene collected from one or several viewpoints, or from a video -- is given. Consequently, they focus on estimates leveraging prior knowledge and measurement by…

Computer Vision and Pattern Recognition · Computer Science 2020-12-17 Erik Gärtner , Aleksis Pirinen , Cristian Sminchisescu

Solving depth estimation with monocular cameras enables the possibility of widespread use of cameras as low-cost depth estimation sensors in applications such as autonomous driving and robotics. However, learning such a scalable depth…

Computer Vision and Pattern Recognition · Computer Science 2020-07-30 Bin Cheng , Inderjot Singh Saggu , Raunak Shah , Gaurav Bansal , Dinesh Bharadia
‹ Prev 1 8 9 10 Next ›