中文
相关论文

相关论文: Towards Ambiguity-Free Spatial Foundation Model: R…

200 篇论文

Monocular depth estimation remains challenging, as foundation models such as Depth Anything V2 (DA-V2) struggle with real-world images that are far from the training distribution. We introduce Re-Depth Anything, a test-time self-supervision…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Ananta R. Bhattarai , Helge Rhodin

Majority of the perception methods in robotics require depth information provided by RGB-D cameras. However, standard 3D sensors fail to capture depth of transparent objects due to refraction and absorption of light. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2021-04-02 Luyang Zhu , Arsalan Mousavian , Yu Xiang , Hammad Mazhar , Jozef van Eenbergen , Shoubhik Debnath , Dieter Fox

Open vocabulary object detection (OVD) aims at seeking an optimal object detector capable of recognizing objects from both base and novel categories. Recent advances leverage knowledge distillation to transfer insightful knowledge from…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Jiaming Li , Jiacheng Zhang , Jichang Li , Ge Li , Si Liu , Liang Lin , Guanbin Li

Estimating depth from a sequence of posed RGB images is a fundamental computer vision task, with applications in augmented reality, path planning etc. Prior work typically makes use of previous frames in a multi view stereo framework,…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Mohamed Sayed , Filippo Aleotti , Jamie Watson , Zawar Qureshi , Guillermo Garcia-Hernando , Gabriel Brostow , Sara Vicente , Michael Firman

While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions typically rely on…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Xianjin Wu , Dingkang Liang , Tianrui Feng , Kui Xia , Yumeng Zhang , Xiaofan Li , Xiao Tan , Xiang Bai

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Haoyu Zhang , Meng Liu , Zaijing Li , Haokun Wen , Weili Guan , Yaowei Wang , Liqiang Nie

We consider the generic problem of detecting low-level structures in images, which includes segmenting the manipulated parts, identifying out-of-focus pixels, separating shadow regions, and detecting concealed objects. Whereas each such…

计算机视觉与模式识别 · 计算机科学 2023-03-22 Weihuang Liu , Xi Shen , Chi-Man Pun , Xiaodong Cun

In this work, we present a panoramic metric depth foundation model that generalizes across diverse scene distances. We explore a data-in-the-loop paradigm from the view of both data construction and framework design. We collect a…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Xin Lin , Meixi Song , Dizhe Zhang , Wenxuan Lu , Haodong Li , Bo Du , Ming-Hsuan Yang , Truong Nguyen , Lu Qi

True video understanding requires making sense of non-lambertian scenes where the color of light arriving at the camera sensor encodes information about not just the last object it collided with, but about multiple mediums -- colored…

计算机视觉与模式识别 · 计算机科学 2019-04-05 Jean-Baptiste Alayrac , João Carreira , Andrew Zisserman

Depth estimation is a core problem in robotic perception and vision tasks, but 3D reconstruction from a single image presents inherent uncertainties. Current depth estimation models primarily rely on inter-image relationships for supervised…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Jinchang Zhang , Guoyu Lu

We introduce ByDeWay, a training-free framework designed to enhance the performance of Multimodal Large Language Models (MLLMs). ByDeWay uses a novel prompting strategy called Layered-Depth-Based Prompting (LDP), which improves spatial…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Rajarshi Roy , Devleena Das , Ankesh Banerjee , Arjya Bhattacharjee , Kousik Dasgupta , Subarna Tripathi

In this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to directly predict absolute depth, recent advancements advocate…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Chi Zhang , Wei Yin , Gang Yu , Zhibin Wang , Tao Chen , Bin Fu , Joey Tianyi Zhou , Chunhua Shen

In the field of monocular depth estimation (MDE), many models with excellent zero-shot performance in general scenes emerge recently. However, these methods often fail in predicting non-Lambertian surfaces, such as transparent or mirror…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Junrui Zhang , Jiaqi Li , Yachuan Huang , Yiran Wang , Jinghong Zheng , Liao Shen , Zhiguo Cao

Recovering the scene depth from a single image is an ill-posed problem that requires additional priors, often referred to as monocular depth cues, to disambiguate different 3D interpretations. In recent works, those priors have been learned…

计算机视觉与模式识别 · 计算机科学 2020-08-18 Lam Huynh , Phong Nguyen-Ha , Jiri Matas , Esa Rahtu , Janne Heikkila

In safety-critical domains, linguistic ambiguity can have severe consequences; a vague command like "Pass me the vial" in a surgical setting could lead to catastrophic errors. Yet, most embodied AI research overlooks this, assuming…

人工智能 · 计算机科学 2026-04-16 Jiayu Ding , Haoran Tang , Hongbo Jin , Wei Gao , Ge Li

Self-supervised multi-frame monocular depth estimation relies on the geometric consistency between successive frames under the assumption of a static scene. However, the presence of moving objects in dynamic scenes introduces inevitable…

计算机视觉与模式识别 · 计算机科学 2024-07-15 Sungmin Woo , Wonjoon Lee , Woo Jin Kim , Dogyoon Lee , Sangyoun Lee

Recently, Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multi-modal context comprehension. However, they still suffer from hallucination problems referring to generating inconsistent outputs with the…

计算机视觉与模式识别 · 计算机科学 2024-09-02 Xiaoye Qu , Jiashuo Sun , Wei Wei , Yu Cheng

Commercial RGB-D cameras often produce noisy, incomplete depth maps for non-Lambertian objects. Traditional depth completion methods struggle to generalize due to the limited diversity and scale of training data. Recent advances exploit…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Wenzhou Lyu , Jialing Lin , Wenqi Ren , Ruihao Xia , Feng Qian , Yang Tang

Depth estimation is of critical interest for scene understanding and accurate 3D reconstruction. Most recent approaches in depth estimation with deep learning exploit geometrical structures of standard sharp images to predict corresponding…

计算机视觉与模式识别 · 计算机科学 2018-09-07 Marcela Carvalho , Bertrand Le Saux , Pauline Trouvé-Peloux , Andrés Almansa , Frédéric Champagnat

Vision-Language-Action (VLA) models have recently achieved remarkable progress in robotic perception and control, yet most existing approaches primarily rely on VLM trained using 2D images, which limits their spatial understanding and…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Zhifeng Rao , Wenlong Chen , Lei Xie , Xia Hua , Dongfu Yin , Zhen Tian , F. Richard Yu