中文
相关论文

相关论文: DA$^{2}$: Depth Anything in Any Direction

200 篇论文

We propose a method for metric-scale monocular depth estimation. Inferring depth from a single image is an ill-posed problem due to the loss of scale from perspective projection during the image formation process. Any scale chosen is a…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Ziyao Zeng , Yangchao Wu , Hyoungseob Park , Daniel Wang , Fengyu Yang , Stefano Soatto , Dong Lao , Byung-Woo Hong , Alex Wong

Vision-and-language navigation (VLN) requires an embodied agent to ground natural-language instructions into executable navigation actions in unseen environments. Existing zero-shot methods typically rely on additional waypoint prediction…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Kai Sheng , Liuyi Wang , Haojie Dai , Jinlong Li , Yongrui Qin , Zongtao He , Chengju Liu , Qijun Chen

Obstacle avoidance in unmanned aerial vehicles (UAVs), as a fundamental capability, has gained increasing attention with the growing focus on spatial intelligence. However, current obstacle-avoidance methods mainly depend on limited…

机器人学 · 计算机科学 2026-03-09 Xiangkai Zhang , Dizhe Zhang , WenZhuo Cao , Zhaoliang Wan , Yingjie Niu , Lu Qi , Xu Yang , Zhiyong Liu

Self-supervised methods have showed promising results on depth estimation task. However, previous methods estimate the target depth map and camera ego-motion simultaneously, underusing multi-frame correlation information and ignoring the…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Songchun Zhang , Chunhui Zhao

Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation,…

音频与语音处理 · 电气工程与系统科学 2025-06-04 Huadai Liu , Tianyi Luo , Kaicheng Luo , Qikai Jiang , Peiwen Sun , Jialei Wang , Rongjie Huang , Qian Chen , Wen Wang , Xiangtai Li , Shiliang Zhang , Zhijie Yan , Zhou Zhao , Wei Xue

We introduce ZeroVO, a novel visual odometry (VO) algorithm that achieves zero-shot generalization across diverse cameras and environments, overcoming limitations in existing methods that depend on predefined or static camera calibration…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Lei Lai , Zekai Yin , Eshed Ohn-Bar

Spherical coordinate systems, which are ubiquitous in astronomy, cannot be shown without distortion on flat, two-dimensional surfaces. This poses challenges for the two complementary phases of visual exploration -- making discoveries in…

天体物理仪器与方法 · 物理学 2018-07-18 C. J. Fluke , D. G. Barnes

Although existing monocular depth estimation methods have made great progress, predicting an accurate absolute depth map from a single image is still challenging due to the limited modeling capacity of networks and the scale ambiguity…

计算机视觉与模式识别 · 计算机科学 2022-10-07 Jie Xiang , Yun Wang , Lifeng An , Haiyang Liu , Zijun Wang , Jian Liu

Deep learning approaches achieve prominent success in 3D semantic segmentation. However, collecting densely annotated real-world 3D datasets is extremely time-consuming and expensive. Training models on synthetic data and generalizing on…

计算机视觉与模式识别 · 计算机科学 2022-07-22 Runyu Ding , Jihan Yang , Li Jiang , Xiaojuan Qi

Omnidirectional scene understanding is vital for various downstream applications, such as embodied AI, autonomous driving, and immersive environments, yet remains challenging due to geometric distortion and complex spatial relations in…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Xinshen Zhang , Tongxi Fu , Xu Zheng

The recent development of \emph{foundation models} for monocular depth estimation such as Depth Anything paved the way to zero-shot monocular depth estimation. Since it returns an affine-invariant disparity map, the favored technique to…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Rémi Marsal , Alexandre Chapoutot , Philippe Xu , David Filliat

This work addresses the task of zero-shot monocular depth estimation. A recent advance in this field has been the idea of utilising Text-to-Image foundation models, such as Stable Diffusion. Foundation models provide a rich and generic…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Denis Zavadski , Damjan Kalšan , Carsten Rother

Omnidirectional Videos (or 360{\deg} videos) are widely used in Virtual Reality (VR) to facilitate immersive and interactive viewing experiences. However, the limited spatial resolution in 360{\deg} videos does not allow for each degree of…

多媒体 · 计算机科学 2025-06-19 Arbind Agrahari Baniya , Tsz-Kwan Lee , Peter W. Eklund , Sunil Aryal

This paper presents a novel preconditioning strategy for the classic 8-point algorithm (8-PA) for estimating an essential matrix from 360-FoV images (i.e., equirectangular images) in spherical projection. To alleviate the effect of uneven…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Bolivar Solarte , Chin-Hsuan Wu , Kuan-Wei Lu , Min Sun , Wei-Chen Chiu , Yi-Hsuan Tsai

Unsupervised domain adaptation methods for panoramic semantic segmentation utilize real pinhole images or low-cost synthetic panoramic images to transfer segmentation models to real panoramic images. However, these methods struggle to…

计算机视觉与模式识别 · 计算机科学 2025-01-08 Jing Jiang , Sicheng Zhao , Jiankun Zhu , Wenbo Tang , Zhaopan Xu , Jidong Yang , Guoping Liu , Tengfei Xing , Pengfei Xu , Hongxun Yao

Time-of-Flight (ToF) cameras possess compact design and high measurement precision to be applied to various robot tasks. However, their limited sensing range restricts deployment in large-scale scenarios. Depth completion has emerged as a…

机器人学 · 计算机科学 2026-03-24 Juncheng Chen , Tiancheng Lai , Xingpeng Wang , Bingxin Liao , Baozhe Zhang , Chao Xu , Yanjun Cao

Underwater infrastructure requires frequent inspection and maintenance due to harsh marine conditions. Current reliance on human divers or remotely operated vehicles is limited by perceptual and operational challenges, especially around…

计算机视觉与模式识别 · 计算机科学 2025-10-30 Hongjie Zhang , Gideon Billings , Stefan B. Williams

One practical approach to infer 3D scene structure from a single image is to retrieve a closely matching 3D model from a database and align it with the object in the image. Existing methods rely on supervised training with images and pose…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Pattaramanee Arsomngern , Sasikarn Khwanmuang , Matthias Nießner , Supasorn Suwajanakorn

This paper designs a technique route to generate high-quality panoramic image with depth information, which involves two critical research hotspots: fusion of LiDAR and image data and image stitching. For the fusion of 3D points and image…

计算机视觉与模式识别 · 计算机科学 2020-10-28 Hao Ma , Jingbin Liu , Zhirong Hu , Hongyu Qiu , Dong Xu , Zemin Wang , Xiaodong Gong , Sheng Yang

Large models have shown generalization across datasets for many low-level vision tasks, like depth estimation, but no such general models exist for scene flow. Even though scene flow has wide potential use, it is not used in practice…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Yiqing Liang , Abhishek Badki , Hang Su , James Tompkin , Orazio Gallo