English
Related papers

Related papers: DINO in the Room: Leveraging 2D Foundation Models …

200 papers

Self-supervised learning (SSL) leverages vast unannotated medical datasets, yet steep technical barriers limit adoption by clinical researchers. We introduce Vision Foundry, a code-free, HIPAA-compliant platform that democratizes…

3D semantic segmentation provides high-level scene understanding for applications in robotics, autonomous systems, \textit{etc}. Traditional methods adapt exclusively to either task-specific goals (open-vocabulary segmentation) or scene…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Doriand Petit , Steve Bourgeois , Vincent Gay-Bellile , Florian Chabot , Loïc Barthe

Light sheet fluorescence microscopy (LSM) enables high-resolution, three-dimensional (3D) imaging of biological specimens, providing rich volumetric data for studying cellular organization, pathology, and vascular networks. However, the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Adina Scheinfeld , Haotan Zhang , Shang Mu , Rudolf L. M. van Herten , Lucas Stoffl , Ali Erturk , Zhuhao Wu , Johannes C. Paetzold

Foundation models leverage large-scale pretraining to capture extensive knowledge, demonstrating generalization in a wide range of language tasks. By comparison, vision foundation models (VFMs) often exhibit uneven improvements across…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Shiqi Huang , Yipei Wang , Natasha Thorley , Alexander Ng , Shaheer Saeed , Mark Emberton , Shonit Punwani , Veeru Kasivisvanathan , Dean Barratt , Daniel Alexander , Yipeng Hu

Medical image segmentation is a crucial method for assisting professionals in diagnosing various diseases through medical imaging. However, various factors such as noise, blurriness, and low contrast often hinder the accurate diagnosis of…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Jeonghyun Noh , Wangsu Jeon , Jinsun Park

Foundation vision models are increasingly adopted in medical image analysis. Due to domain shift, these pretrained models misalign with medical image segmentation needs without being fully fine-tuned or lightly adapted. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zhuonan Liang , Wei Guo , Jie Gan , Yaxuan Song , Runnan Chen , Hang Chang , Weidong Cai

We address the problem of extending the capabilities of vision foundation models such as DINO, SAM, and CLIP, to 3D tasks. Specifically, we introduce a novel method to uplift 2D image features into Gaussian Splatting representations of 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Juliette Marrie , Romain Menegaux , Michael Arbel , Diane Larlus , Julien Mairal

Multimodal medical image fusion plays an instrumental role in several areas of medical image processing, particularly in disease recognition and tumor detection. Traditional fusion methods tend to process each modality independently before…

Image and Video Processing · Electrical Eng. & Systems 2023-10-11 Lin Liu , Xinxin Fan , Chulong Zhang , Jingjing Dai , Yaoqin Xie , Xiaokun Liang

Volumetric models have become a popular representation for 3D scenes in recent years. One of the breakthroughs leading to their popularity was KinectFusion, where the focus is on 3D reconstruction using RGB-D sensors. However, monocular…

Computer Vision and Pattern Recognition · Computer Science 2014-10-27 Victor Adrian Prisacariu , Olaf Kähler , Ming Ming Cheng , Carl Yuheng Ren , Julien Valentin , Philip H. S. Torr , Ian D. Reid , David W. Murray

Foundation models (FMs) have achieved remarkable success across a wide range of applications, from image classification to natural langurage processing, but pose significant challenges for deployment at edge. This has sparked growing…

Machine Learning · Computer Science 2025-07-17 Muhammad Azlan Qazi , Alexandros Iosifidis , Qi Zhang

Real-time object detection has achieved substantial progress through meticulously designed architectures and optimization strategies. However, the pursuit of high-speed inference via lightweight network designs often leads to degraded…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Zijun Liao , Yian Zhao , Xin Shan , Yu Yan , Chang Liu , Lei Lu , Xiangyang Ji , Jie Chen

Embodied tasks require the agent to fully understand 3D scenes simultaneously with its exploration, so an online, real-time, fine-grained and highly-generalized 3D perception model is desperately needed. Since high-quality 3D data is…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Xiuwei Xu , Huangxing Chen , Linqing Zhao , Ziwei Wang , Jie Zhou , Jiwen Lu

Recently, feature upsampling has gained increasing attention owing to its effectiveness in enhancing vision foundation models (VFMs) for pixel-level understanding tasks. Existing methods typically rely on high-resolution features from the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xiaoqiong Liu , Heng Fan

Advances in fluorescence microscopy enable acquisition of 3D image volumes with better image quality and deeper penetration into tissue. Segmentation is a required step to characterize and analyze biological structures in the images and…

Computer Vision and Pattern Recognition · Computer Science 2019-04-16 Chichen Fu , Soonam Lee , David Joon Ho , Shuo Han , Paul Salama , Kenneth W. Dunn , Edward J. Delp

Remarkable progress in 2D Vision-Language Models (VLMs) has spurred interest in extending them to 3D settings for tasks like 3D Question Answering, Dense Captioning, and Visual Grounding. Unlike 2D VLMs that typically process images through…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Haoyuan Li , Yanpeng Zhou , Yufei Gao , Tao Tang , Jianhua Han , Yujie Yuan , Dave Zhenyu Chen , Jiawang Bian , Hang Xu , Xiaodan Liang

Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent space, previous work typically adapts the CLIP encoder to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xinwei He , Yansong Zheng , Qianru Han , Zhichuan Wang , Yuxuan Cai , Yang Zhou , Jingbo Xia , Yulong Wang , Jinhai Xiang , Xiang Bai

Grounding natural language to the physical world is a ubiquitous topic with a wide range of applications in computer vision and robotics. Recently, 2D vision-language models such as CLIP have been widely popularized, due to their impressive…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Georgios Tziafas , Yucheng Xu , Zhibin Li , Hamidreza Kasaei

In order to navigate complex traffic environments, self-driving vehicles must recognize many semantic classes pertaining to vulnerable road users or traffic control devices. However, many safety-critical objects (e.g., construction worker)…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Anqi Joyce Yang , James Tu , Nikita Dvornik , Enxu Li , Raquel Urtasun

LiDAR-camera 3D representation pretraining has shown significant promise for 3D perception tasks and related applications. However, two issues widely exist in this framework: 1) Solely keyframes are used for training. For example, in…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Tianfang Sun , Zhizhong Zhang , Xin Tan , Yanyun Qu , Yuan Xie

The application of 3D ViTs to medical image segmentation has seen remarkable strides, somewhat overshadowing the budding advancements in Convolutional Neural Network (CNN)-based models. Large kernel depthwise convolution has emerged as a…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Ho Hin Lee , Quan Liu , Qi Yang , Xin Yu , Shunxing Bao , Yuankai Huo , Bennett A. Landman
‹ Prev 1 4 5 6 7 8 10 Next ›