中文
相关论文

相关论文: Hybrid Structure-from-Motion and Camera Relocaliza…

200 篇论文

Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and large intra-class appearance variations. Objects may leave and re-enter the field of…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Dmytro Klepachevskyi , Alexander Wong , Sirisha Rambhatla , Yuhao Chen

We present EgoTAP, a heatmap-to-3D pose lifting method for highly accurate stereo egocentric 3D pose estimation. Severe self-occlusion and out-of-view limbs in egocentric camera views make accurate pose estimation a challenging problem. To…

计算机视觉与模式识别 · 计算机科学 2024-02-29 Taeho Kang , Youngki Lee

In the last twenty years, Structure from Motion (SfM) has been a constant research hotspot in the fields of photogrammetry, computer vision, robotics etc., whereas real-time performance is just a recent topic of growing interest. This work…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Zongqian Zhan , Yifei Yu , Rui Xia , Wentian Gan , Hong Xie , Giulio Perda , Luca Morelli , Fabio Remondino , Xin Wang

Understanding ego-motion and surrounding vehicle state is essential to enable automated driving and advanced driving assistance technologies. Typical approaches to solve this problem use fusion of multiple sensors such as LiDAR, camera, and…

计算机视觉与模式识别 · 计算机科学 2020-05-07 Jun Hayakawa , Behzad Dariush

Camera geo-localization from a monocular video is a fundamental task for video analysis and autonomous navigation. Although 3D reconstruction is a key technique to obtain camera poses, monocular 3D reconstruction in a large environment…

计算机视觉与模式识别 · 计算机科学 2018-08-28 Kazuya Iwami , Satoshi Ikehata , Kiyoharu Aizawa

The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Daniele Materia , Francesco Ragusa , Giovanni Maria Farinella

This paper introduces a preprocessing technique to speed up Structure-from-Motion (SfM) based pose estimation, which is critical for real-time applications like augmented reality (AR), virtual reality (VR), and robotics. Our method…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Joji Joseph , Bharadwaj Amrutur , Shalabh Bhatnagar

Finding local features that are repeatable across multiple views is a cornerstone of sparse 3D reconstruction. The classical image matching paradigm detects keypoints per-image once and for all, which can yield poorly-localized features and…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Philipp Lindenberger , Paul-Edouard Sarlin , Viktor Larsson , Marc Pollefeys

We consider the problem of transferring a temporal action segmentation system initially designed for exocentric (fixed) cameras to an egocentric scenario, where wearable cameras capture video data. The conventional supervised approach…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Camillo Quattrocchi , Antonino Furnari , Daniele Di Mauro , Mario Valerio Giuffrida , Giovanni Maria Farinella

Visual localization algorithms, i.e., methods that estimate the camera pose of a query image in a known scene, are core components of many applications, including self-driving cars and augmented / mixed reality systems. State-of-the-art…

计算机视觉与模式识别 · 计算机科学 2025-04-25 Vojtech Panek , Qunjie Zhou , Yaqing Ding , Sérgio Agostinho , Zuzana Kukelova , Torsten Sattler , Laura Leal-Taixé

Many monocular visual SLAM algorithms are derived from incremental structure-from-motion (SfM) methods. This work proposes a novel monocular SLAM method which integrates recent advances made in global SfM. In particular, we present two main…

计算机视觉与模式识别 · 计算机科学 2017-10-20 Chengzhou Tang , Oliver Wang , Ping Tan

Predicting turn-taking in multiparty conversations has many practical applications in human-computer/robot interaction. However, the complexity of human communication makes it a challenging task. Recent advances have shown that synchronous…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Mehdi Fatan , Emanuele Mincato , Dimitra Pintzou , Mariella Dimiccoli

Lately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. However, the adaptation of these models to egocentric videos has…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Anna Kukleva , Fadime Sener , Edoardo Remelli , Bugra Tekin , Eric Sauser , Bernt Schiele , Shugao Ma

In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise localization of the interval within an untrimmed egocentric…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Yisen Feng , Haoyu Zhang , Qiaohui Chu , Meng Liu , Weili Guan , Yaowei Wang , Liqiang Nie

We propose a complete pipeline that allows object detection and simultaneously estimate the pose of these multiple object instances using just a single image. A novel "keypoint regression" scheme with a cross-ratio term is introduced that…

计算机视觉与模式识别 · 计算机科学 2018-09-28 Ankit Dhall

Accurately estimating the 3D pose of the camera wearer in egocentric video sequences is crucial to modeling human behavior in virtual and augmented reality applications. The task presents unique challenges due to the limited visibility of…

计算机视觉与模式识别 · 计算机科学 2024-11-08 Luca Scofano , Alessio Sampieri , Edoardo De Matteis , Indro Spinelli , Fabio Galasso

Video-language pre-training (VLP) has become increasingly important due to its ability to generalize to various vision and language tasks. However, existing egocentric VLP frameworks utilize separate video and language encoders and learn…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Shraman Pramanick , Yale Song , Sayan Nag , Kevin Qinghong Lin , Hardik Shah , Mike Zheng Shou , Rama Chellappa , Pengchuan Zhang

We leverage unsupervised learning of depth, egomotion, and camera intrinsics to improve the performance of single-image semantic segmentation, by enforcing 3D-geometric and temporal consistency of segmentation masks across video frames. The…

计算机视觉与模式识别 · 计算机科学 2020-05-22 Ankita Pasad , Ariel Gordon , Tsung-Yi Lin , Anelia Angelova

Reconstruction of 3D neural fields from posed images has emerged as a promising method for self-supervised representation learning. The key challenge preventing the deployment of these 3D scene learners on large-scale video data is their…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Cameron Smith , Yilun Du , Ayush Tewari , Vincent Sitzmann

As the number of installed cameras grows, so do the compute resources required to process and analyze all the images captured by these cameras. Video analytics enables new use cases, such as smart cities or autonomous driving. At the same…

计算机视觉与模式识别 · 计算机科学 2021-12-01 Daniel Rivas , Francesc Guim , Jordà Polo , David Carrera
‹ 上一页 1 8 9 10 下一页 ›