English
Related papers

Related papers: VisionFM: a Multi-Modal Multi-Task Vision Foundati…

200 papers

In recent years, large language models (LLMs) have demonstrated remarkable potential across various medical applications. Building on this foundation, multimodal large language models (MLLMs) integrate LLMs with visual models to process…

Computation and Language · Computer Science 2025-03-11 Xiaoyi Liang , Mouxiao Bian , Moxin Chen , Lihao Liu , Junjun He , Jie Xu , Lin Li

Medical large vision-language models (LVLMs) have demonstrated promising performance across various single-image question answering (QA) benchmarks, yet their capability in processing multi-image clinical scenarios remains underexplored.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Xikai Yang , Juzheng Miao , Yuchen Yuan , Jiaze Wang , Qi Dou , Jinpeng Li , Pheng-Ann Heng

Vision foundation models (VFMs) are pre-trained on extensive image datasets to learn general representations for diverse types of data. These models can subsequently be fine-tuned for specific downstream tasks, significantly boosting…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Shansong Wang , Mojtaba Safari , Qiang Li , Chih-Wei Chang , Richard LJ Qiu , Justin Roper , David S. Yu , Xiaofeng Yang

With the rapid advancement of deep learning, particularly in the field of medical image analysis, an increasing number of Vision-Language Models (VLMs) are being widely applied to solve complex health and biomedical challenges. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Haiyang Yu , Siyang Yi , Ke Niu , Minghan Zhuo , Bin Li

Visual Foundation Models (VFMs) are becoming ubiquitous in computer vision, powering systems for diverse tasks such as object detection, image classification, segmentation, pose estimation, and motion tracking. VFMs are capitalizing on…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Sandeep Gupta , Roberto Passerone

Recent advances in AI combine large language models (LLMs) with vision encoders that bring forward unprecedented technical capabilities to leverage for a wide range of healthcare applications. Focusing on the domain of radiology,…

Current RF machine-learning pipelines rely on task-specific deep networks for modulation classification and related tasks, but these models require custom architectures and labeled datasets for each problem, generalize poorly across channel…

Signal Processing · Electrical Eng. & Systems 2026-02-17 Hang Zou , Bohao Wang , Yu Tian , Lina Bariah , Chongwen Huang , Samson Lasaulce , Mérouane Debbah

Retinal disease diagnosis is critical in preventing vision loss and reducing socioeconomic burdens. Globally, over 2.2 billion people are affected by some form of vision impairment, resulting in annual productivity losses estimated at $411…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Youssef Sabiri , Walid Houmaidi , Amine Abouaomar

Foundation models (FMs) are large-scale deep learning models trained on massive datasets, often using self-supervised learning techniques. These models serve as a versatile base for a wide range of downstream tasks, including those in…

Machine Learning · Computer Science 2025-01-17 Wasif Khan , Seowung Leem , Kyle B. See , Joshua K. Wong , Shaoting Zhang , Ruogu Fang

We introduce UViM, a unified approach capable of modeling a wide range of computer vision tasks. In contrast to previous models, UViM has the same functional form for all tasks; it requires no task-specific modifications which require…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Alexander Kolesnikov , André Susano Pinto , Lucas Beyer , Xiaohua Zhai , Jeremiah Harmsen , Neil Houlsby

Vision systems to see and reason about the compositional nature of visual scenes are fundamental to understanding our world. The complex relations between objects and their locations, ambiguities, and variations in the real-world…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Muhammad Awais , Muzammal Naseer , Salman Khan , Rao Muhammad Anwer , Hisham Cholakkal , Mubarak Shah , Ming-Hsuan Yang , Fahad Shahbaz Khan

The pervasive uncertainty and dynamic nature of real-world environments present significant challenges for the widespread implementation of machine-driven Intelligent Decision-Making (IDM) systems. Consequently, IDM should possess the…

Artificial Intelligence · Computer Science 2023-05-17 Ying Wen , Ziyu Wan , Ming Zhou , Shufang Hou , Zhe Cao , Chenyang Le , Jingxiao Chen , Zheng Tian , Weinan Zhang , Jun Wang

Background: RETFound, a self-supervised, retina-specific foundation model (FM), showed potential in downstream applications. However, its comparative performance with traditional deep learning (DL) models remains incompletely understood.…

Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable promise in generating visually grounded responses. However, their application in the medical domain is hindered by unique challenges. For instance, most VLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Lingxiao Luo , Bingda Tang , Xuanzhong Chen , Rong Han , Ting Chen

Segmenting 3D blood vessels is a critical yet challenging task in medical image analysis. This is due to significant imaging modality-specific variations in artifacts, vascular patterns and scales, signal-to-noise ratios, and background…

Image and Video Processing · Electrical Eng. & Systems 2025-03-18 Bastian Wittmann , Yannick Wattenberg , Tamaz Amiranashvili , Suprosanna Shit , Bjoern Menze

Background: Cardiovascular diseases (CVDs) are the leading cause of death globally. The use of artificial intelligence (AI) methods - in particular, deep learning (DL) - has been on the rise lately for the analysis of different CVD-related…

In ophthalmology, early fundus screening is an economic and effective way to prevent blindness caused by ophthalmic diseases. Clinically, due to the lack of medical resources, manual diagnosis is time-consuming and may delay the condition.…

Computer Vision and Pattern Recognition · Computer Science 2021-02-17 Ning Li , Tao Li , Chunyu Hu , Kai Wang , Hong Kang

Modern medical records include a vast amount of multimodal free text clinical data and imaging data from radiology, cardiology, and digital pathology. Fully mining such big data requires multitasking; otherwise, occult but important aspects…

Image and Video Processing · Electrical Eng. & Systems 2024-04-25 Chuang Niu , Qing Lyu , Christopher D. Carothers , Parisa Kaviani , Josh Tan , Pingkun Yan , Mannudeep K. Kalra , Christopher T. Whitlow , Ge Wang

Civil aviation is a cornerstone of global transportation and commerce, and ensuring its safety, efficiency and customer satisfaction is paramount. Yet conventional Artificial Intelligence (AI) solutions in aviation remain siloed and narrow,…

Artificial Intelligence · Computer Science 2026-01-19 Wenbin Li , Jingling Wu , Xiaoyong Lin. Jing Chen , Cong Chen

Ultrasound imaging is widely used in clinical diagnosis due to its non-invasive nature and real-time capabilities. However, traditional ultrasound diagnostics relies heavily on physician expertise and is often hampered by suboptimal image…

Image and Video Processing · Electrical Eng. & Systems 2025-12-18 Yuncheng Jiang , Chun-Mei Feng , Jinke Ren , Jun Wei , Zixun Zhang , Yiwen Hu , Yunbi Liu , Rui Sun , Xuemei Tang , Juan Du , Xiang Wan , Yong Xu , Bo Du , Xin Gao , Guangyu Wang , Shaohua Zhou , Shuguang Cui , Zhen Li