中文
相关论文

相关论文: Formula-Supervised Visual-Geometric Pre-training

200 篇论文

We propose a novel fine-grained cross-view localization method that estimates the 3 Degrees of Freedom pose of a ground-level image in an aerial image of the surroundings by matching fine-grained features between the two images. The pose is…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Zimin Xia , Alexandre Alahi

The significance of cross-view 3D geometric modeling capabilities for autonomous driving is self-evident, yet existing Vision-Language Models (VLMs) inherently lack this capability, resulting in their mediocre performance. While some…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Jie Wang , Guang Li , Zhijian Huang , Chenxu Dang , Hangjun Ye , Yahong Han , Long Chen

Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly,…

Point cloud registration is a fundamental task in 3D vision. Most existing methods only use geometric information for registration. Recently proposed RGB-D registration methods primarily focus on feature fusion or improving feature…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Congjia Chen , Shen Yan , Yufu Qu

In this work, we aim to improve the 3D reasoning ability of Transformers in multi-view 3D human pose estimation. Recent works have focused on end-to-end learning-based transformer designs, which struggle to resolve geometric information…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Ziwei Liao , Jialiang Zhu , Chunyu Wang , Han Hu , Steven L. Waslander

Formula-driven supervised learning (FDSL) has been shown to be an effective method for pre-training vision transformers, where ExFractalDB-21k was shown to exceed the pre-training effect of ImageNet-21k. These studies also indicate that…

计算机视觉与模式识别 · 计算机科学 2023-03-03 Sora Takashima , Ryo Hayamizu , Nakamasa Inoue , Hirokatsu Kataoka , Rio Yokota

Nowadays, robotics, AR, and 3D modeling applications attract considerable attention to single-view depth estimation (SVDE) as it allows estimating scene geometry from a single RGB image. Recent works have demonstrated that the accuracy of…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Nikolay Patakin , Mikhail Romanov , Anna Vorontsova , Mikhail Artemyev , Anton Konushin

Existing approaches to vision-language pre-training (VLP) heavily rely on an object detector based on bounding boxes (regions), where salient objects are first detected from images and then a Transformer-based model is used for cross-modal…

多媒体 · 计算机科学 2021-08-24 Ming Yan , Haiyang Xu , Chenliang Li , Bin Bi , Junfeng Tian , Min Gui , Wei Wang

Scene understanding based on 3D Gaussian Splatting (3DGS) has recently achieved notable advances. Although 3DGS related methods have efficient rendering capabilities, they fail to address the inherent contradiction between the anisotropic…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Q. G. Duan , Benyun Zhao , Mingqiao Han Yijun Huang , Ben M. Chen

3D Visual Grounding (3DVG) seeks to locate target objects in 3D scenes using natural language descriptions, enabling downstream applications such as augmented reality and robotics. Existing approaches typically rely on labeled 3D data and…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Rong Li , Shijie Li , Lingdong Kong , Xulei Yang , Junwei Liang

Grasping unknown objects from a single view has remained a challenging topic in robotics due to the uncertainty of partial observation. Recent advances in large-scale models have led to benchmark solutions such as GraspNet-1Billion.…

机器人学 · 计算机科学 2025-07-17 Hao Chen , Takuya Kiyokawa , Zhengtao Hu , Weiwei Wan , Kensuke Harada

Reconstructing 3D models from 2D images is one of the fundamental problems in computer vision. In this work, we propose a deep learning technique for 3D object reconstruction from a single image. Contrary to recent works that either use 3D…

计算机视觉与模式识别 · 计算机科学 2020-05-06 K L Navaneet , Ansu Mathew , Shashank Kashyap , Wei-Chih Hung , Varun Jampani , R. Venkatesh Babu

Advanced visual localization techniques encompass image retrieval challenges and 6 Degree-of-Freedom (DoF) camera pose estimation, such as hierarchical localization. Thus, they must extract global and local features from input images.…

计算机视觉与模式识别 · 计算机科学 2022-12-27 Wenzheng Song , Ran Yan , Boshu Lei , Takayuki Okatani

We study to generate novel views of indoor scenes given sparse input views. The challenge is to achieve both photorealism and view consistency. We present SparseGNV: a learning framework that incorporates 3D structures and image generative…

计算机视觉与模式识别 · 计算机科学 2023-05-12 Weihao Cheng , Yan-Pei Cao , Ying Shan

We introduce FaceGPT, a self-supervised learning framework for Large Vision-Language Models (VLMs) to reason about 3D human faces from images and text. Typical 3D face reconstruction methods are specialized algorithms that lack semantic…

计算机视觉与模式识别 · 计算机科学 2024-06-12 Haoran Wang , Mohit Mendiratta , Christian Theobalt , Adam Kortylewski

In this work, we address the challenging task of 3D object recognition without the reliance on real-world 3D labeled data. Our goal is to predict the 3D shape, size, and 6D pose of objects within a single RGB-D image, operating at the…

计算机视觉与模式识别 · 计算机科学 2023-10-20 Mayank Lunayach , Sergey Zakharov , Dian Chen , Rares Ambrus , Zsolt Kira , Muhammad Zubair Irshad

Deep neural network models have achieved remarkable progress in 3D scene understanding while trained in the closed-set setting and with full labels. However, the major bottleneck is that these models do not have the capacity to recognize…

计算机视觉与模式识别 · 计算机科学 2025-02-20 Kangcheng Liu , Yong-Jin Liu , Baoquan Chen

Fine-grained visual classification (FGVC) is much more challenging than traditional classification tasks due to the inherently subtle intra-class object variations. Recent works mainly tackle this problem by focusing on how to locate the…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Ruoyi Du , Dongliang Chang , Ayan Kumar Bhunia , Jiyang Xie , Zhanyu Ma , Yi-Zhe Song , Jun Guo

3D semantic occupancy prediction has become a crucial perception task for comprehensive scene understanding in autonomous driving. While recent advances have explored 3D Gaussian splatting for occupancy modeling to substantially reduce…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Xiaoyang Yan , Muleilan Pei , Shaojie Shen

Feed-forward 3D reconstruction models, represented by Visual Geometry Grounded Transformer (VGGT), jointly predict multiple visual geometry tasks such as depth estimation, camera pose prediction, and point cloud reconstruction in a single…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Yipu Zhang , Jintao Cheng , Weilun Feng , Jiehao Luo , Chuanguang Yang , Zhulin An , Yongjun Xu , Wei Zhang