English
Related papers

Related papers: V$^2$-SfMLearner: Learning Monocular Depth and Ego…

200 papers

Event cameras are neuromorphically inspired sensors that sparsely and asynchronously report brightness changes. Their unique characteristics of high temporal resolution, high dynamic range, and low power consumption make them well-suited…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Haitao Meng , Chonghao Zhong , Sheng Tang , Lian JunJia , Wenwei Lin , Zhenshan Bing , Yi Chang , Gang Chen , Alois Knoll

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

Robotics · Computer Science 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Robust and accurate trajectory estimation of mobile agents such as people and robots is a key requirement for providing spatial awareness for emerging capabilities such as augmented reality or autonomous interaction. Although currently…

This work presents Depth Anything V2. Without pursuing fancy techniques, we aim to reveal crucial findings to pave the way towards building a powerful monocular depth estimation model. Notably, compared with V1, this version produces much…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Lihe Yang , Bingyi Kang , Zilong Huang , Zhen Zhao , Xiaogang Xu , Jiashi Feng , Hengshuang Zhao

In the realm of modern diagnostic technology, video capsule endoscopy (VCE) is a standout for its high efficacy and non-invasive nature in diagnosing various gastrointestinal (GI) conditions, including obscure bleeding. Importantly, for the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Hechen Li , Yanan Wu , Long Bai , An Wang , Tong Chen , Hongliang Ren

Despite significant progress in Vision-Language Pre-training (VLP), current approaches predominantly emphasize feature extraction and cross-modal comprehension, with limited attention to generating or transforming visual content. This gap…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Ziyang Zhang , Yang Yu , Yucheng Chen , Xulei Yang , Si Yong Yeo

Single-image depth estimation is essential for endoscopy tasks such as localization, reconstruction, and augmented reality. Most existing methods in surgical scenes focus on in-domain depth estimation, limiting their real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Qingyao Tian , Zhen Chen , Huai Liao , Xinyan Huang , Lujie Li , Sebastien Ourselin , Hongbin Liu

Monocular depth estimation (MDE) plays a pivotal role in various computer vision applications, such as robotics, augmented reality, and autonomous driving. Despite recent advancements, existing methods often fail to meet key requirements…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Andrii Litvynchuk , Ivan Livinsky , Anand Ravi , Nima Kalantari , Andrii Tsarov

Visual-inertial odometry (VIO) is the pose estimation backbone for most AR/VR and autonomous robotic systems today, in both academia and industry. However, these systems are highly sensitive to the initialization of key parameters such as…

Degenerative spinal pathologies are highly prevalent among the elderly population. Timely diagnosis of osteoporotic fractures and other degenerative deformities facilitates proactive measures to mitigate the risk of severe back pain and…

Image and Video Processing · Electrical Eng. & Systems 2023-12-11 Hellena Hempe , Alexander Bigalke , Mattias P. Heinrich

The escalating global mortality and morbidity rates associated with gastrointestinal (GI) bleeding, compounded by the complexities and limitations of traditional endoscopic methods, underscore the urgent need for a critical review of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Tanisha Singh , Shreshtha Jha , Nidhi Bhatt , Palak Handa , Nidhi Goel , Sreedevi Indu

Reliable and real-time 3D reconstruction and localization functionality is a crucial prerequisite for the navigation of actively controlled capsule endoscopic robots as an emerging, minimally invasive diagnostic and therapeutic technology…

Accurate monocular metric depth estimation (MMDE) is crucial to solving downstream tasks in 3D perception and modeling. However, the remarkable accuracy of recent MMDE methods is confined to their training domains. These methods fail to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Luigi Piccinelli , Christos Sakaridis , Yung-Hsu Yang , Mattia Segu , Siyuan Li , Wim Abbeloos , Luc Van Gool

Endoscopic surgery is the gold standard for robotic-assisted minimally invasive surgery, offering significant advantages in early disease detection and precise interventions. However, the complexity of surgical scenes, characterized by high…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Guankun Wang , Rui Tang , Mengya Xu , Long Bai , Huxin Gao , Hongliang Ren

Unsupervised deep learning is a promising method in brain MRI registration to reduce the reliance on anatomical labels, while still achieving anatomically accurate transformations. For the Learn2Reg2024 LUMIR challenge, we propose…

Image and Video Processing · Electrical Eng. & Systems 2024-12-31 Lukas Förner , Kartikay Tehlan , Thomas Wendler

Real-time ego-motion tracking for endoscope is a significant task for efficient navigation and robotic automation of endoscopy. In this paper, a novel framework is proposed to perform real-time ego-motion tracking for endoscope. Firstly, a…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Liangjing Shao , Benshuang Chen , Shuting Zhao , Xinrong Chen

Flexible endoscopes for colonoscopy present several limitations due to their inherent complexity, resulting in patient discomfort and lack of intuitiveness for clinicians. Robotic devices together with autonomous control represent a viable…

Monocular depth inference has gained tremendous attention from researchers in recent years and remains as a promising replacement for expensive time-of-flight sensors, but issues with scale acquisition and implementation overhead still…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Kenny Chen , Alexandra Pogue , Brett T. Lopez , Ali-akbar Agha-mohammadi , Ankur Mehta

Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Shuyao Shi , Kang G. Shin

Self-supervised monocular depth estimation methods have been increasingly given much attention due to the benefit of not requiring large, labelled datasets. Such self-supervised methods require high-quality salient features and consequently…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Xiaotong Guo , Huijie Zhao , Shuwei Shao , Xudong Li , Baochang Zhang
‹ Prev 1 4 5 6 7 8 10 Next ›