English
Related papers

Related papers: Harnessing Foundation Models for Robust and Genera…

200 papers

Robust, high-precision global localization is fundamental to a wide range of outdoor robotics applications. Conventional fusion methods use low-accuracy pseudorange based GNSS measurements ($>>5m$ errors) and can only yield a coarse…

The monocular visual-inertial system (VINS), which consists one camera and one low-cost inertial measurement unit (IMU), is a popular approach to achieve accurate 6-DOF state estimation. However, such locally accurate visual-inertial…

Computer Vision and Pattern Recognition · Computer Science 2018-03-06 Tong Qin , Perliang Li , Shaojie Shen

Generalizable dense feature matching in endoscopic images is crucial for robot-assisted tasks, including 3D reconstruction, navigation, and surgical scene understanding. Yet, it remains a challenge due to difficult visual conditions (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Bingyu Yang , Qingyao Tian , Yimeng Geng , Huai Liao , Xinyan Huang , Jiebo Luo , Hongbin Liu

This paper addresses the problem of estimating the 3-DoF camera pose for a ground-level image with respect to a satellite image that encompasses the local surroundings. We propose a novel end-to-end approach that leverages the learning of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Zhenbo Song , Xianghui Ze , Jianfeng Lu , Yujiao Shi

Vision-language models, such as CLIP, have achieved significant success in aligning visual and textual representations, becoming essential components of many multi-modal large language models (MLLMs) like LLaVA and OpenFlamingo. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Shizhan Gong , Yankai Jiang , Qi Dou , Farzan Farnia

Endoscopic surgery relies on two-dimensional views, posing challenges for surgeons in depth perception and instrument manipulation. While Monocular Visual Simultaneous Localization and Mapping (MVSLAM) has emerged as a promising solution,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 G. Manni , C. Lauretti , F. Prata , R. Papalia , L. Zollo , P. Soda

Despite significant algorithmic advances in vision-based positioning, a comprehensive probabilistic framework to study its performance has remained unexplored. The main objective of this paper is to develop such a framework using ideas from…

Information Theory · Computer Science 2024-09-17 Haozhou Hu , Harpreet S. Dhillon , R. Michael Buehrer

Previous evaluations on 6DoF object pose tracking have presented obvious limitations along with the development of this area. In particular, the evaluation protocols are not unified for different methods, the widely-used YCBV dataset…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Yang Li , Fan Zhong , Xin Wang , Shuangbing Song , Jiachen Li , Xueying Qin , Changhe Tu

In this report, we present the first place solution to the ECCV 2024 BRAVO Challenge, where a model is trained on Cityscapes and its robustness is evaluated on several out-of-distribution datasets. Our solution leverages the powerful…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Tommie Kerssies , Daan de Geus , Gijs Dubbelman

Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends leverage Vision-Language Models (VLMs) for semantic supervision, these multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Mika Feng , Pierre Gallin-Martel , Koichi Ito , Takafumi Aoki

Visual servo based on traditional image matching methods often requires accurate keypoint correspondence for high precision control. However, keypoint detection or matching tends to fail in challenging scenarios with inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Anzhe Chen , Hongxiang Yu , Shuxin Li , Yuxi Chen , Zhongxiang Zhou , Wentao Sun , Rong Xiong , Yue Wang

We present a robust deep learning based 6 degrees-of-freedom (DoF) localization system for endoscopic capsule robots. Our system mainly focuses on localization of endoscopic capsule robots inside the GI tract using only visual information…

Computer Vision and Pattern Recognition · Computer Science 2017-05-17 Mehmet Turan , Yasin Almalioglu , Ender Konukoglu , Metin Sitti

Learning whole-body mobile manipulation via imitation is essential for generalizing robotic skills to diverse environments and complex tasks. However, this goal is hindered by significant challenges, particularly in effectively processing…

Robotics · Computer Science 2025-09-29 Yue Su , Chubin Zhang , Sijin Chen , Liufan Tan , Yansong Tang , Jianan Wang , Xihui Liu

We focus on the generalization ability of the 6-DoF grasp detection method in this paper. While learning-based grasp detection methods can predict grasp poses for unseen objects using the grasp distribution learned from the training set,…

Robotics · Computer Science 2024-04-03 Haoxiang Ma , Modi Shi , Boyang Gao , Di Huang

In medical and industrial domains, providing guidance for assembly processes can be critical to ensure efficiency and safety. Errors in assembly can lead to significant consequences such as extended surgery times and prolonged manufacturing…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Hannah Schieber , Shiyu Li , Niklas Corell , Philipp Beckerle , Julian Kreimeier , Daniel Roth

Visual localization is the task of estimating a 6-DoF camera pose of a query image within a provided 3D reference map. Thanks to recent advances in various 3D sensors, 3D point clouds are becoming a more accurate and affordable option for…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Minjung Kim , Junseo Koo , Gunhee Kim

Large-scale vision foundation models have demonstrated remarkable success across various tasks, underscoring their robust generalization capabilities. While their proficiency in two-view correspondence has been explored, their effectiveness…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Görkay Aydemir , Weidi Xie , Fatma Güney

We propose a framework for extracting the bone surface from B-mode images employing the eigenspace minimum variance (ESMV) beamformer and a ridge detection method. We show that an ESMV beamformer with a rank-1 signal subspace can preserve…

Medical Physics · Physics 2016-09-07 Saeed Mehdizadeh , Sebastien Muller , Gabriel Kiss , Tonni F. Johansen , Sverre Holm

While recent advances in neural radiance field enable realistic digitization for large-scale scenes, the image-capturing process is still time-consuming and labor-intensive. Previous works attempt to automate this process using the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Xiao Chen , Quanyi Li , Tai Wang , Tianfan Xue , Jiangmiao Pang

In endoscopic procedures, autonomous tracking of abnormal regions and following circumferential cutting markers can significantly reduce the cognitive burden on endoscopists. However, conventional model-based pipelines are fragile for each…

Robotics · Computer Science 2025-08-21 Chi Kit Ng , Long Bai , Guankun Wang , Yupeng Wang , Huxin Gao , Kun Yuan , Chenhan Jin , Tieyong Zeng , Hongliang Ren