English
Related papers

Related papers: VisionFM: a Multi-Modal Multi-Task Vision Foundati…

200 papers

Foundation models (FMs) have emerged as a transformative paradigm in medical image analysis, offering the potential to provide generalizable, task-agnostic solutions across a wide range of clinical tasks and imaging modalities. Their…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Karma Phuntsho , Abdullah , Kyungmi Lee , Ickjai Lee , Euijoon Ahn

Medical vision-and-language models (MVLMs) have attracted substantial interest due to their capability to offer a natural language interface for interpreting complex medical data. Their applications are versatile and have the potential to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Qi Chen , Ruoshan Zhao , Sinuo Wang , Vu Minh Hieu Phan , Anton van den Hengel , Johan Verjans , Zhibin Liao , Minh-Son To , Yong Xia , Jian Chen , Yutong Xie , Qi Wu

Vision-Language Foundation Models (VLFMs) have made remarkable progress on various multimodal tasks, such as image captioning, image-text retrieval, visual question answering, and visual grounding. However, most methods rely on training…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Yue Zhou , Zhihang Zhong , Xue Yang

Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 You Qin , Kai Liu , Shengqiong Wu , Kai Wang , Shijian Deng , Yapeng Tian , Junbin Xiao , Yazhou Xing , Yinghao Ma , Bobo Li , Roger Zimmermann , Lei Cui , Furu Wei , Jiebo Luo , Hao Fei

Wireless foundation models (WFMs) have recently demonstrated promising capabilities, jointly performing multiple wireless functions and adapting effectively to new environments. However, while current WFMs process only one modality,…

Signal Processing · Electrical Eng. & Systems 2026-02-20 Ahmed Aboulfotouh , Hatem Abou-Zeid

Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individual tasks. Emerging research indicates that large…

Artificial Intelligence · Computer Science 2024-11-06 Dawei Dai , Xu Long , Li Yutang , Zhang Yuanhui , Shuyin Xia

Foundation vision-language models are currently transforming computer vision, and are on the rise in medical imaging fueled by their very promising generalization capabilities. However, the initial attempts to transfer this new paradigm to…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Julio Silva-Rodríguez , Hadi Chakor , Riadh Kobbi , Jose Dolz , Ismail Ben Ayed

The advent of foundation models, which are pre-trained on vast datasets, has ushered in a new era of computer vision, characterized by their robustness and remarkable zero-shot generalization capabilities. Mirroring the transformative…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Xu Liu , Tong Zhou , Yuanxin Wang , Yuping Wang , Qinjingwen Cao , Weizhi Du , Yonghuan Yang , Junjun He , Yu Qiao , Yiqing Shen

Breast cancer is one of the leading causes of death among women worldwide. We introduce Mammo-FM, the first foundation model specifically for mammography, pretrained on the largest and most diverse dataset to date - 140,677 patients…

Large Vision-Language Models offer a new paradigm for AI-driven image understanding, enabling models to perform tasks without task-specific training. This flexibility holds particular promise across medicine, where expert-annotated data is…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Anita Rau , Mark Endo , Josiah Aklilu , Jaewoo Heo , Khaled Saab , Alberto Paderno , Jeffrey Jopling , F. Christopher Holsinger , Serena Yeung-Levy

Technological advances in artificial intelligence (AI) have enabled the development of large vision language models (LVLMs) that are trained on millions of paired image and text samples. Subsequent research efforts have demonstrated great…

Computation and Language · Computer Science 2024-11-28 Kalina P. Slavkova , Melanie Traughber , Oliver Chen , Robert Bakos , Shayna Goldstein , Dan Harms , Bradley J. Erickson , Khan M. Siddiqui

Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments following language instructions. Despite significant progress in large Vision-Language…

This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to general-purpose…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Chunyuan Li , Zhe Gan , Zhengyuan Yang , Jianwei Yang , Linjie Li , Lijuan Wang , Jianfeng Gao

We envision the "virtual eye" as a next-generation, AI-powered platform that uses interconnected foundation models to simulate the eye's intricate structure and biological function across all scales. Advances in AI, imaging, and multiomics…

Tissues and Organs · Quantitative Biology 2025-05-12 Yue Wu , Yibo Guo , Yulong Yan , Jiancheng Yang , Xin Zhou , Ching-Yu Cheng , Danli Shi , Mingguang He

Multimodal learning, especially large-scale multimodal pre-training, has developed rapidly over the past few years and led to the greatest advances in artificial intelligence (AI). Despite its effectiveness, understanding the underlying…

Neural and Evolutionary Computing · Computer Science 2022-08-18 Haoyu Lu , Qiongyi Zhou , Nanyi Fei , Zhiwu Lu , Mingyu Ding , Jingyuan Wen , Changde Du , Xin Zhao , Hao Sun , Huiguang He , Ji-Rong Wen

Retinal foundation models aim to learn generalizable representations from diverse retinal images, facilitating label-efficient model adaptation across various ophthalmic tasks. Despite their success, current retinal foundation models are…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Kai Yu , Yang Zhou , Yang Bai , Zhi Da Soh , Xinxing Xu , Rick Siow Mong Goh , Ching-Yu Cheng , Yong Liu

There is substantial interest in developing artificial intelligence systems to support radiologists across tasks ranging from segmentation to report generation. Existing computed tomography (CT) foundation models have largely focused on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Rubén Moreno-Aguado , Alba Magallón , Victor Moreno , Yingying Fang , Guang Yang

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Jingzhi Li , Changjiang Luo , Ruoyu Chen , Hua Zhang , Wenqi Ren , Jianhou Gan , Xiaochun Cao

Advances in artificial intelligence (AI) have achieved expert-level performance in medical imaging applications. Notably, self-supervised vision-language foundation models can detect a broad spectrum of pathologies without relying on…

Computers and Society · Computer Science 2024-02-23 Yuzhe Yang , Yujia Liu , Xin Liu , Avanti Gulhane , Domenico Mastrodicasa , Wei Wu , Edward J Wang , Dushyant W Sahani , Shwetak Patel
‹ Prev 1 4 5 6 7 8 10 Next ›