中文
相关论文

相关论文: On the Coupling of Depth and Egomotion Networks fo…

200 篇论文

Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Shuyao Shi , Kang G. Shin

Recent supervised multi-view depth estimation networks have achieved promising results. Similar to all supervised approaches, these networks require ground-truth data during training. However, collecting a large amount of multi-view depth…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Jiayu Yang , Jose M. Alvarez , Miaomiao Liu

We present a model for the joint estimation of disparity and motion. The model is based on learning about the interrelations between images from multiple cameras, multiple frames in a video, or the combination of both. We show that learning…

计算机视觉与模式识别 · 计算机科学 2013-12-17 Kishore Konda , Roland Memisevic

We present GLNet, a self-supervised framework for learning depth, optical flow, camera pose and intrinsic parameters from monocular video - addressing the difficulty of acquiring realistic ground-truth for such tasks. We propose three…

计算机视觉与模式识别 · 计算机科学 2019-09-10 Yuhua Chen , Cordelia Schmid , Cristian Sminchisescu

We integrate two powerful ideas, geometry and deep visual representation learning, into recurrent network architectures for mobile visual scene understanding. The proposed networks learn to "lift" and integrate 2D visual features over time…

计算机视觉与模式识别 · 计算机科学 2019-04-10 Hsiao-Yu Fish Tung , Ricson Cheng , Katerina Fragkiadaki

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans actively coordinate…

机器人学 · 计算机科学 2026-03-11 Justin Yu , Yide Shentu , Di Wu , Pieter Abbeel , Ken Goldberg , Philipp Wu

In this work we present a self-supervised learning framework to simultaneously train two Convolutional Neural Networks (CNNs) to predict depth and surface normals from a single image. In contrast to most existing frameworks which represent…

计算机视觉与模式识别 · 计算机科学 2019-03-04 Huangying Zhan , Chamara Saroj Weerasekera , Ravi Garg , Ian Reid

Person image generation aims to perform non-rigid deformation on source images, which generally requires unaligned data pairs for training. Recently, self-supervised methods express great prospects in this task by merging the disentangled…

计算机视觉与模式识别 · 计算机科学 2022-12-15 Zijian Wang , Xingqun Qi , Kun Yuan , Muyi Sun

Current Structure-from-Motion (SfM) methods typically follow a two-stage pipeline, combining learned or geometric pairwise reasoning with a subsequent global optimization step. In contrast, we propose a data-driven multi-view reasoning…

计算机视觉与模式识别 · 计算机科学 2025-05-09 Qitao Zhao , Amy Lin , Jeff Tan , Jason Y. Zhang , Deva Ramanan , Shubham Tulsiani

View-graph is an essential input to large-scale structure from motion (SfM) pipelines. Accuracy and efficiency of large-scale SfM is crucially dependent on the input view-graph. Inconsistent or inaccurate edges can lead to inferior or wrong…

计算机视觉与模式识别 · 计算机科学 2017-12-05 Rajvi Shah , Visesh Chari , P J Narayanan

Multi-frame methods improve monocular depth estimation over single-frame approaches by aggregating spatial-temporal information via feature matching. However, the spatial-temporal feature leads to accuracy degradation in dynamic scenes. To…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Jiquan Zhong , Xiaolin Huang , Xiao Yu

Self-supervised monocular scene flow estimation, aiming to understand both 3D structures and 3D motions from two temporally consecutive monocular images, has received increasing attention for its simple and economical sensor setup. However,…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Zijie Jiang , Masatoshi Okutomi

Self-supervised learning is the key to unlocking generic computer vision systems. By eliminating the reliance on ground-truth annotations, it allows scaling to much larger data quantities. Unfortunately, self-supervised monocular depth…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Jaime Spencer , Chris Russell , Simon Hadfield , Richard Bowden

In this work we introduce a new self-supervised, semi-parametric approach for synthesizing novel views of a vehicle starting from a single monocular image. Differently from parametric (i.e. entirely learning-based) methods, we show how…

计算机视觉与模式识别 · 计算机科学 2020-10-13 Andrea Palazzi , Luca Bergamini , Simone Calderara , Rita Cucchiara

Depth estimation from a single image represents a very exciting challenge in computer vision. While other image-based depth sensing techniques leverage on the geometry between different viewpoints (e.g., stereo or structure from motion),…

计算机视觉与模式识别 · 计算机科学 2018-10-29 Pierluigi Zama Ramirez , Matteo Poggi , Fabio Tosi , Stefano Mattoccia , Luigi Di Stefano

Current end-to-end autonomous driving methods typically learn only from expert planning data collected from a single ego vehicle, severely limiting the diversity of learnable driving policies and scenarios. However, a critical yet…

计算机视觉与模式识别 · 计算机科学 2025-07-16 Lin Liu , Caiyan Jia , Ziying Song , Hongyu Pan , Bencheng Liao , Wenchao Sun , Yongchang Zhang , Lei Yang , Yandan Luo

Federated learning (FL) based magnetic resonance (MR) image reconstruction can facilitate learning valuable priors from multi-site institutions without violating patient's privacy for accelerating MR imaging. However, existing methods rely…

图像与视频处理 · 电气工程与系统科学 2023-05-11 Juan Zou , Cheng Li , Ruoyou Wu , Tingrui Pei , Hairong Zheng , Shanshan Wang

Multi-frame depth estimation improves over single-frame approaches by also leveraging geometric relationships between images via feature matching, in addition to learning appearance-based features. In this paper we revisit feature matching…

计算机视觉与模式识别 · 计算机科学 2022-06-14 Vitor Guizilini , Rares Ambrus , Dian Chen , Sergey Zakharov , Adrien Gaidon

Training deep neural networks to estimate the viewpoint of objects requires large labeled training datasets. However, manually labeling viewpoints is notoriously hard, error-prone, and time-consuming. On the other hand, it is relatively…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Siva Karthik Mustikovela , Varun Jampani , Shalini De Mello , Sifei Liu , Umar Iqbal , Carsten Rother , Jan Kautz

Multi-view depth estimation plays a critical role in reconstructing and understanding the 3D world. Recent learning-based methods have made significant progress in it. However, multi-view depth estimation is fundamentally a…

计算机视觉与模式识别 · 计算机科学 2022-05-06 Kai Cheng , Hao Chen , Wei Yin , Guangkai Xu , Xuejin Chen