大尺寸注意力下的自监督单目深度估计
计算机视觉与模式识别
2024-09-27 v1
摘要
自监督单目深度估计因不依赖标记训练数据而成为一种有前景的 method。大多数 method 结合卷积和 Transformer 来建模长程依赖,以准确估计深度。然而,Transformer 将 2D 图像特征视为 1D 序列,位置编码虽能一定程度缓解不同特征块之间空间信息的损失,却倾向于忽视通道特征,这限制了深度估计的性能。本文提出一种自监督单目深度估计网络以获取更细致的细节。具体而言,我们提出一种基于大尺寸注意力的解码器,能够在不损失二维特征结构的前提下建模长程依赖,同时保持特征通道的适应性。此外,我们引入一个上采样模块,以准确恢复深度图中的细节。我们的方法在 KITTI 数据集上取得了竞争性的结果。
引用
@article{arxiv.2409.17895,
title = {Self-supervised Monocular Depth Estimation with Large Kernel Attention},
author = {Xuezhi Xiang and Yao Wang and Lei Zhang and Denis Ombati and Himaloy Himu and Xiantong Zhen},
journal= {arXiv preprint arXiv:2409.17895},
year = {2024}
}
备注
The paper is under consideration at 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)