English
Related papers

Related papers: Leveraging Generic Foundation Models for Multimoda…

200 papers

This paper explores feature prediction as a stand-alone objective for unsupervised learning from video and introduces V-JEPA, a collection of vision models trained solely using a feature prediction objective, without the use of pretrained…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Adrien Bardes , Quentin Garrido , Jean Ponce , Xinlei Chen , Michael Rabbat , Yann LeCun , Mahmoud Assran , Nicolas Ballas

The integration of deep learning systems into healthcare has been hindered by the resource-intensive process of data annotation and the inability of these systems to generalize to different data distributions. Foundation models, which are…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Mohammed Baharoon , Waseem Qureshi , Jiahong Ouyang , Yanwu Xu , Abdulrhman Aljouie , Wei Peng

As the potential of foundation models in visual tasks has garnered significant attention, pretraining these models before downstream tasks has become a crucial step. The three key factors in pretraining foundation models are the pretraining…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Keumgang Cha , Junghoon Seo , Taekyung Lee

Foundation models have shown outstanding performance and generalization capabilities across domains. Since most studies on foundation models mainly focus on the pretraining phase, a naive strategy to minimize a single task-specific loss is…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Dohwan Ko , Joonmyung Choi , Hyeong Kyu Choi , Kyoung-Woon On , Byungseok Roh , Hyunwoo J. Kim

Joint Embedding Predictive Architectures (JEPA) have emerged as a powerful framework for learning general-purpose representations. However, these models often lack interpretability and suffer from inefficiencies due to dense embedding…

Machine Learning · Computer Science 2025-04-24 Max Hartman , Lav Varshney

Although deep learning have revolutionized abdominal multi-organ segmentation, models often struggle with generalization due to training on small, specific datasets. With the recent emergence of large-scale datasets, some important…

Image and Video Processing · Electrical Eng. & Systems 2025-02-20 Ziyan Huang , Zhongying Deng , Jin Ye , Haoyu Wang , Yanzhou Su , Tianbin Li , Hui Sun , Junlong Cheng , Jianpin Chen , Junjun He , Yun Gu , Shaoting Zhang , Lixu Gu , Yu Qiao

Estimation of pain intensity from facial expressions captured in videos has an immense potential for health care applications. Given the challenges related to subjective variations of facial expressions, and operational capture conditions,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 R. Gnana Praveen , Eric Granger , Patrick Cardinal

Foundation models like CLIP allow zero-shot transfer on various tasks without additional training data. Yet, the zero-shot performance is less competitive than a fully supervised one. Thus, to enhance the performance, fine-tuning and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Beier Zhu , Kaihua Tang , Qianru Sun , Hanwang Zhang

Foundational models are trained on extensive datasets to capture the general trends of a domain. However, in medical imaging, the scarcity of data makes pre-training for every domain, modality, or task challenging. Instead of building…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Mohammad Areeb Qazi , Munachiso S Nwadike , Ibrahim Almakky , Mohammad Yaqub , Numan Saeed

Federated Learning (FL) enables decentralized model training across multiple parties while preserving privacy. However, most FL systems assume clients hold only unimodal data, limiting their real-world applicability, as institutions often…

Machine Learning · Computer Science 2025-04-17 Yu Zhang , Qingfeng Du , Jiaqi Lv

Foundation models provide robust embeddings for diverse tasks, including medical imaging. We evaluate embeddings from seven general and medical-specific foundation models (e.g., DenseNet121, BiomedCLIP, MedImageInsight, Rad-DINO,…

Large pretrained transformers are increasingly being developed as generalised foundation models which can underpin powerful task-specific artificial intelligence models. Histopathology foundation models show great promise across many tasks,…

Image and Video Processing · Electrical Eng. & Systems 2024-09-10 Jack Breen , Katie Allen , Kieran Zucker , Lucy Godson , Nicolas M. Orsi , Nishant Ravikumar

A unified foundation model for medical time series -- pretrained on open access and ethics board-approved medical corpora -- offers the potential to reduce annotation burdens, minimize model customization, and enable robust transfer across…

Machine Learning · Computer Science 2025-12-17 Hao Li , Bowen Deng , Chang Xu , Zhiyuan Feng , Viktor Schlegel , Yu-Hao Huang , Yizheng Sun , Jingyuan Sun , Kailai Yang , Yiyao Yu , Jiang Bian

Multimodal survival prediction, a crucial yet challenging task, demands the integration of multimodal medical data (\eg Whole Slide Images (WSIs) and Genomic Profiles) to achieve accurate prognostic modeling. Given the inherent…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Huayi Wang , Haochao Ying , Yuyang Xu , Qiyao Zheng , jun wang , Cheng Zhang , Ying Sun , Jian Wu

The dominant paradigm for end-to-end robot learning focuses on optimizing task-specific objectives that solve a single robotic problem such as picking up an object or reaching a target position. However, recent work on high-capacity models…

Robotics · Computer Science 2024-01-02 Samuel Schmidgall , Ji Woong Kim , Alan Kuntz , Ahmed Ezzat Ghazi , Axel Krieger

Recent advancements in multimodal foundation models have showcased impressive capabilities in understanding and reasoning with visual and textual information. Adapting these foundation models trained for general usage to specialized domains…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Hejie Cui , Lingjun Mao , Xin Liang , Jieyu Zhang , Hui Ren , Quanzheng Li , Xiang Li , Carl Yang

We present MeFEm, a vision model based on a modified Joint Embedding Predictive Architecture (JEPA) for biometric and medical analysis from facial images. Key modifications include an axial stripe masking strategy to focus learning on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Yury Borets , Stepan Botman

Radiological analysis increasingly benefits from pretrained visual representations that can support heterogeneous downstream tasks across imaging modalities. In this work, we introduce OmniRad, a self-supervised radiological foundation…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Luca Zedda , Andrea Loddo , Cecilia Di Ruberto

With the emergence of various molecular tasks and massive datasets, how to perform efficient training has become an urgent yet under-explored issue in the area. Data pruning (DP), as an oft-stated approach to saving training burdens,…

Machine Learning · Computer Science 2024-09-04 Dingshuo Chen , Zhixun Li , Yuyan Ni , Guibin Zhang , Ding Wang , Qiang Liu , Shu Wu , Jeffrey Xu Yu , Liang Wang

Fine-grained glomerular subtyping is central to kidney biopsy interpretation, but clinically valuable labels are scarce and difficult to obtain. Existing computational pathology approaches instead tend to evaluate coarse diseased…