English
Related papers

Related papers: Generalist Foundation Models from a Multimodal Dat…

200 papers

Deep networks trained on millions of facial images are believed to be closely approaching human-level performance in face recognition. However, open world face recognition still remains a challenge. Although, 3D face recognition has an…

Computer Vision and Pattern Recognition · Computer Science 2020-12-03 Syed Zulqarnain Gilani , Ajmal Mian

The construction of three-dimensional multi-modal tissue maps provides an opportunity to spur interdisciplinary innovations across temporal and spatial scales through information integration. While the preponderance of effort is allocated…

Large vision-language models like CLIP are increasingly used in medical imaging tasks due to their ability to align images and text without the need for extensive labeled data. This makes them particularly useful for applications like image…

Machine Learning · Computer Science 2025-12-22 Jasmine Vu , Shivanand Sheshappanavar

Novel Coronavirus disease (COVID-19) is a highly contagious respiratory infection that has had devastating effects on the world. Recently, new COVID-19 variants are emerging making the situation more challenging and threatening. Evaluation…

With the widespread application of artificial intelligence (AI), particularly deep learning (DL) and vision large language models (VLLMs), in skin disease diagnosis, the need for interpretability becomes crucial. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Yuhao Shen , Liyuan Sun , Yan Xu , Wenbin Liu , Shuping Zhang , Shawn Afvari , Zhongyi Han , Jiaoyan Song , Yongzhi Ji , Tao Lu , Xiaonan He , Xin Gao , Juexiao Zhou

Generating accurate and clinically meaningful radiology reports from chest X-ray images remains a significant challenge in medical AI. While recent vision-language models achieve strong results in general radiology report generation, they…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Nikolay Nechaev , Evgeniia Przhezdzetskaia , Dmitry Umerenkov , Dmitry V. Dylov

Clinicians spend a significant amount of time reviewing medical images and transcribing their findings regarding patient diagnosis, referral and treatment in text form. Vision-language models (VLMs), which automatically interpret images and…

Existing vision-text contrastive learning like CLIP aims to match the paired image and caption embeddings while pushing others apart, which improves representation transferability and supports zero-shot prediction. However, medical…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Zifeng Wang , Zhenbang Wu , Dinesh Agarwal , Jimeng Sun

Cardiovascular diseases (CVD) remain a leading health concern and contribute significantly to global mortality rates. While clinical advancements have led to a decline in CVD mortality, accurately identifying individuals who could benefit…

Image and Video Processing · Electrical Eng. & Systems 2024-11-18 Minfeng Xu , Chen-Chen Fan , Yan-Jie Zhou , Wenchao Guo , Pan Liu , Jing Qi , Le Lu , Hanqing Chao , Kunlun He

The emergence of Large Language Models (LLMs) presents unprecedented opportunities to revolutionize medical contrastive vision-language pre-training. In this paper, we show how LLMs can facilitate large-scale supervised pre-training,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Yingtai Li , Haoran Lai , Xiaoqian Zhou , Shuai Ming , Wenxin Ma , Wei Wei , Shaohua Kevin Zhou

Computed Tomography (CT) is a frequently utilized imaging technology that is employed in the clinical diagnosis of many disorders. However, clinical diagnosis, data storage, and management are posed huge challenges by a huge volume of…

Image and Video Processing · Electrical Eng. & Systems 2024-05-02 Siyi Xun , Qiaoyu Li , Xiaohong Liu , Guangtao Zhai , Mingxiang Wu , Tao Tan

Medical report generation has achieved remarkable advancements yet has still been faced with several challenges. First, the inherent imbalance in the distribution of normal and abnormal cases may lead models to exhibit a biased focus on…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Zhixuan Chen , Luyang Luo , Yequan Bie , Hao Chen

The advancement of artificial intelligence (AI) for organ segmentation and tumor detection is propelled by the growing availability of computed tomography (CT) datasets with detailed, per-voxel annotations. However, these AI models often…

Image and Video Processing · Electrical Eng. & Systems 2024-05-29 Jie Liu , Yixiao Zhang , Kang Wang , Mehmet Can Yavuz , Xiaoxi Chen , Yixuan Yuan , Haoliang Li , Yang Yang , Alan Yuille , Yucheng Tang , Zongwei Zhou

The COVID-19 pandemic has spread globally for several months. Because its transmissibility and high pathogenicity seriously threaten people's lives, it is crucial to accurately and quickly detect COVID-19 infection. Many recent studies have…

Image and Video Processing · Electrical Eng. & Systems 2021-02-15 Xin He , Shihao Wang , Xiaowen Chu , Shaohuai Shi , Jiangping Tang , Xin Liu , Chenggang Yan , Jiyong Zhang , Guiguang Ding

AI-driven models have demonstrated significant potential in automating radiology report generation for chest X-rays. However, there is no standardized benchmark for objectively evaluating their performance. To address this, we present…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Xiaoman Zhang , Hong-Yu Zhou , Xiaoli Yang , Oishi Banerjee , Julián N. Acosta , Josh Miller , Ouwen Huang , Pranav Rajpurkar

Vision-language models have been extensively explored across a wide range of tasks, achieving satisfactory performance; however, their application in medical imaging remains underexplored. In this work, we propose a unified framework -…

Image and Video Processing · Electrical Eng. & Systems 2024-07-18 Khai Le-Duc , Ryan Zhang , Ngoc Son Nguyen , Tan-Hanh Pham , Anh Dao , Ba Hung Ngo , Anh Totti Nguyen , Truong-Son Hy

The development of successful artificial intelligence models for chest X-ray analysis relies on large, diverse datasets with high-quality annotations. While several databases of chest X-ray images have been released, most include disease…

Image and Video Processing · Electrical Eng. & Systems 2024-05-21 Nicolás Gaggion , Candelaria Mosquera , Lucas Mansilla , Julia Mariel Saidman , Martina Aineseder , Diego H. Milone , Enzo Ferrante

Large language models (LLMs), such as ChatGPT, have demonstrated impressive capabilities in various tasks and attracted an increasing interest as a natural language interface across many domains. Recently, large vision-language models…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Zhihao Chen , Bin Hu , Chuang Niu , Tao Chen , Yuxin Li , Hongming Shan , Ge Wang

Over 1.4 billion chest X-rays (CXRs) are performed annually due to their cost-effectiveness as an initial diagnostic test. This scale of radiological studies provides a significant opportunity to streamline CXR interpretation and…

As the demand for more descriptive machine learning models grows within medical imaging, bottlenecks due to data paucity will exacerbate. Thus, collecting enough large-scale data will require automated tools to harvest data/label pairs from…

Image and Video Processing · Electrical Eng. & Systems 2019-10-01 Bo Zhou , Adam P. Harrison , Jiawen Yao , Chi-Tung Cheng , Jing Xiao , Chien-Hung Liao , Le Lu