English
Related papers

Related papers: DINO in the Room: Leveraging 2D Foundation Models …

200 papers

Although large-scale visual foundation models (VFMs) achieve remarkable performance in semantic understanding, they still underperform in instance-aware dense prediction tasks. They exhibit different biases in representation: for instance,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yachan Guo , JoseLuis Gomez Zurita , Danna Xue , Yi Xiao , AntonioManuel Lopez Pena

Foundation models for interactive segmentation in 2D natural images and videos have sparked significant interest in building 3D foundation models for medical imaging. However, the domain gaps and clinical use cases for 3D medical imaging…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Yufan He , Pengfei Guo , Yucheng Tang , Andriy Myronenko , Vishwesh Nath , Ziyue Xu , Dong Yang , Can Zhao , Benjamin Simon , Mason Belue , Stephanie Harmon , Baris Turkbey , Daguang Xu , Wenqi Li

Although vision foundation models (VFMs) are increasingly reused for biomedical image analysis, it remains unclear whether the latent representations they provide are general enough to support effective transfer and reuse across…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Caterina Fuster-Barceló , Virginie Uhlmann

Diffusion models have achieved significant success in both natural image and medical image domains, encompassing a wide range of applications. Previous investigations in medical images have often been constrained to specific anatomical…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yongrui Yu , Yannian Gu , Shaoting Zhang , Xiaofan Zhang

Indoor scene understanding remains a fundamental challenge in robotics, with direct implications for downstream tasks such as navigation and manipulation. Traditional approaches often rely on closed-set recognition or loop closure, limiting…

Robotics · Computer Science 2025-06-10 Hongming Chen , Yiyang Lin , Ziliang Li , Biyu Ye , Yuying Zhang , Ximin Lyu

Vision Foundation Models (VFMs) have become the cornerstone of modern computer vision, offering robust representations across a wide array of tasks. While recent advances allow these models to handle varying input sizes during training,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Bocheng Zou , Mu Cai , Mark Stanley , Dingfu Lu , Yong Jae Lee

The lifting of 3D structure and camera from 2D landmarks is at the cornerstone of the entire discipline of computer vision. Traditional methods have been confined to specific rigid objects, such as those in Perspective-n-Point (PnP)…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Mosam Dabhi , Laszlo A. Jeni , Simon Lucey

Large-scale pre-trained models, such as Vision Foundation Models (VFMs), have demonstrated impressive performance across various downstream tasks by transferring generalized knowledge, especially when target data is limited. However, their…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Pengchen Liang , Haishan Huang , Bin Pu , Jianguo Chen , Xiang Hua , Jing Zhang , Weibo Ma , Zhuangzhuang Chen , Yiwei Li , Qing Chang

Foundation models (FMs) have shown transformative potential in radiology by performing diverse, complex tasks across imaging modalities. Here, we developed CT-FM, a large-scale 3D image-based pre-trained model designed explicitly for…

Image and Video Processing · Electrical Eng. & Systems 2025-02-27 Suraj Pai , Ibrahim Hadzic , Dennis Bontempi , Keno Bressem , Benjamin H. Kann , Andriy Fedorov , Raymond H. Mak , Hugo J. W. L. Aerts

The 3D point cloud representation plays a crucial role in preserving the geometric fidelity of the physical world, enabling more accurate complex 3D environments. While humans naturally comprehend the intricate relationships between objects…

Computer Vision and Pattern Recognition · Computer Science 2025-01-31 Vishal Thengane , Xiatian Zhu , Salim Bouzerdoum , Son Lam Phung , Yunpeng Li

Neural networks achieve state-of-the-art performance in many supervised learning tasks when the training data distribution matches the test data distribution. However, their performance drops significantly under domain (covariate) shift, a…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Kerem Cekmeceli , Meva Himmetoglu , Guney I. Tombak , Anna Susmelj , Ertunc Erdil , Ender Konukoglu

Vision foundation models (VFMs) are pre-trained on extensive image datasets to learn general representations for diverse types of data. These models can subsequently be fine-tuned for specific downstream tasks, significantly boosting…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Shansong Wang , Mojtaba Safari , Qiang Li , Chih-Wei Chang , Richard LJ Qiu , Justin Roper , David S. Yu , Xiaofeng Yang

Online zero-shot 3D instance segmentation of a progressively reconstructed scene is both a critical and challenging task for embodied applications. With the success of visual foundation models (VFMs) in the image domain, leveraging 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Yijie Tang , Jiazhao Zhang , Yuqing Lan , Yulan Guo , Dezun Dong , Chenyang Zhu , Kai Xu

Unsupervised domain adaptation (UDA) is vital for alleviating the workload of labeling 3D point cloud data and mitigating the absence of labels when facing a newly defined domain. Various methods of utilizing images to enhance the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Jingyi Xu , Weidong Yang , Lingdong Kong , Youquan Liu , Rui Zhang , Qingyuan Zhou , Ben Fei

There has been a debate on whether to use 2D or 3D deep neural networks for volumetric organ segmentation. Both 2D and 3D models have their advantages and disadvantages. In this paper, we present an alternative framework, which trains 2D…

Computer Vision and Pattern Recognition · Computer Science 2018-06-12 Yingda Xia , Lingxi Xie , Fengze Liu , Zhuotun Zhu , Elliot K. Fishman , Alan L. Yuille

Recent advancements in vision foundation models (VFMs) have opened up new possibilities for versatile and efficient visual perception. In this work, we introduce Seal, a novel framework that harnesses VFMs for segmenting diverse automotive…

Computer Vision and Pattern Recognition · Computer Science 2023-10-25 Youquan Liu , Lingdong Kong , Jun Cen , Runnan Chen , Wenwei Zhang , Liang Pan , Kai Chen , Ziwei Liu

Contrastive image-to-LiDAR knowledge transfer, commonly used for learning 3D representations with synchronized images and point clouds, often faces a self-conflict dilemma. This issue arises as contrastive losses unintentionally dissociate…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Yifan Zhang , Junhui Hou

This paper proposes novel methods to enhance the performance of monocular 3D object detection models by leveraging the generalized feature extraction capabilities of a vision foundation model. Unlike traditional CNN-based approaches, which…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Jihyeok Kim , Seongwoo Moon , Sungwon Nah , David Hyunchul Shim

Due to its cost-effectiveness and widespread availability, monocular 3D object detection, which relies solely on a single camera during inference, holds significant importance across various applications, including autonomous driving and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Bonan Ding , Jin Xie , Jing Nie , Jiale Cao , Xuelong Li , Yanwei Pang

Image segmentation is a fundamental task in computer vision aimed at delineating object boundaries within images. Traditional approaches, such as edge detection and variational methods, have been widely explored, while recent advances in…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Junchao Zhou