English
Related papers

Related papers: Generalist Foundation Models from a Multimodal Dat…

200 papers

Computed tomography (CT) plays an important role in lung malignancy diagnostics and therapy assessment and facilitating precision medicine delivery. However, the use of personalized imaging protocols poses a challenge in large-scale…

Image and Video Processing · Electrical Eng. & Systems 2020-04-06 Md Selim , Jie Zhang , Baowei Fei , Guo-Qiang Zhang , Jin Chen

General-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation models. However, these…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Fan Liu , Delong Chen , Zhangqingyun Guan , Xiaocong Zhou , Jiale Zhu , Qiaolin Ye , Liyong Fu , Jun Zhou

An increasing number of public datasets have shown a marked impact on automated organ segmentation and tumor detection. However, due to the small size and partially labeled problem of each dataset, as well as a limited investigation of…

Image and Video Processing · Electrical Eng. & Systems 2024-06-27 Jie Liu , Yixiao Zhang , Jie-Neng Chen , Junfei Xiao , Yongyi Lu , Bennett A. Landman , Yixuan Yuan , Alan Yuille , Yucheng Tang , Zongwei Zhou

Automated radiology report generation from 3D CT volumes often suffers from incomplete pathology coverage. We provide empirical evidence that this limitation stems from a representational bottleneck: contrastive 3D CT embeddings encode…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Renjie Liang , Yiling Ma , Yang Xing , Zhengkang Fan , Jinqian Pan , Chengkun Sun , Li Li , Kuang Gong , Jie Xu

Vision-language foundation models, represented by Contrastive Language-Image Pre-training (CLIP), have gained increasing attention for jointly understanding both vision and textual tasks. However, existing approaches primarily focus on…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Bowen Shi , Peisen Zhao , Zichen Wang , Yuhang Zhang , Yaoming Wang , Jin Li , Wenrui Dai , Junni Zou , Hongkai Xiong , Qi Tian , Xiaopeng Zhang

This study demonstrates the first in-hospital adaptation of a cloud-based AI, similar to ChatGPT, into a secure model for analyzing radiology reports, prioritizing patient data privacy. By employing a unique sentence-level knowledge…

Artificial Intelligence · Computer Science 2024-10-30 Kyungsu Kim , Junhyun Park , Saul Langarica , Adham Mahmoud Alkhadrawi , Synho Do

Supervised crowd counting relies heavily on costly manual labeling, which is difficult and expensive, especially in dense scenes. To alleviate the problem, we propose a novel unsupervised framework for crowd counting, named CrowdCLIP. The…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Dingkang Liang , Jiahao Xie , Zhikang Zou , Xiaoqing Ye , Wei Xu , Xiang Bai

AI-assisted imaging made substantial advances in tumor diagnosis and management. However, a major barrier to developing robust oncology foundation models is the scarcity of large-scale, high-quality annotated datasets, which are limited by…

The coronavirus disease 2019 (COVID-19) pandemic continues to have a tremendous impact on patients and healthcare systems around the world. In the fight against this novel disease, there is a pressing need for rapid and effective screening…

Image and Video Processing · Electrical Eng. & Systems 2020-09-14 Hayden Gunraj , Linda Wang , Alexander Wong

The burgeoning integration of 3D medical imaging into healthcare has led to a substantial increase in the workload of medical professionals. To assist clinicians in their diagnostic processes and alleviate their workload, the development of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Yinda Chen , Che Liu , Xiaoyu Liu , Rossella Arcucci , Zhiwei Xiong

Vision-language foundation models have emerged as powerful general-purpose representation learners with strong potential for multimodal understanding, but their deterministic embeddings often fail to provide the reliability required for…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Ahmad Elallaf , Yu Zhang , Yuktha Priya Masupalli , Jeong Yang , Young Lee , Zechun Cao , Gongbo Liang

The rapid increase of computed tomography (CT) scans and their time-consuming manual analysis have created an urgent need for robust automated analysis techniques in clinical settings. These aim to assist radiologists and help them managing…

Image and Video Processing · Electrical Eng. & Systems 2026-02-24 Theo Di Piazza , Carole Lazarus , Olivier Nempont , Loic Boussel

There is an urgent need for triage and classification of high-volume medical imaging modalities such as computed tomography (CT), which can improve patient care and mitigate radiologist burnout. Study-level CT triage requires calibrated…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Lavsen Dahal , Yubraj Bhandari , Geoffrey D. Rubin , Joseph Y. Lo

Weakly supervised disease classification of CT imaging suffers from poor localization owing to case-level annotations, where even a positive scan can hold hundreds to thousands of negative slices along multiple planes. Furthermore, although…

Computer Vision and Pattern Recognition · Computer Science 2020-11-03 Anindo Saha , Fakrul I. Tushar , Khrystyna Faryna , Vincent M. D'Anniballe , Rui Hou , Maciej A. Mazurowski , Geoffrey D. Rubin , Joseph Y. Lo

Deep models, such as convolutional neural networks (CNNs) and vision transformer (ViT), demonstrate remarkable performance in image classification. However, those deep models require large data to fine-tune, which is impractical in the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Yihang Wu , Muhammad Owais , Reem Kateb , Ahmad Chaddad

Echocardiography records ultrasound videos of the heart, enabling clinicians to assess cardiac function. Recent advances in large-scale vision-language models (VLMs) have spurred interest in automating echocardiographic interpretation.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Ryo Takizawa , Satoshi Kodera , Tempei Kabayama , Ryo Matsuoka , Yuta Ando , Yuto Nakamura , Haruki Settai , Norihiko Takeda

Generative AI has advanced rapidly in medical report generation; however, its application to oral and maxillofacial CBCT reporting remains limited, largely because of the scarcity of high-quality paired CBCT-report data and the intrinsic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Qinxin Wu , Fucheng Niu , Hengchuan Zhu , Yifan Sun , Ye Shen , Xu Li , Han Wu , Leqi Liu , Zhiwen Pan , Zuozhu Liu , Fudong Zhu , Bin Feng

Contrastive Language-Image Pre-training (CLIP) has demonstrated strong generalization across a wide range of visual tasks by leveraging large-scale English-image pairs. However, its extension to low-resource languages remains limited due to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Dahyun Chung , Donghyun Shin , Yujin Sung , Seunggi Moon , Jinwoo Jeon , Byung-Jun Lee

Recent advancements in Computer Assisted Diagnosis have shown promising performance in medical imaging tasks, particularly in chest X-ray analysis. However, the interaction between these models and radiologists has been primarily limited to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Yunsoo Kim , Jinge Wu , Yusuf Abdulle , Yue Gao , Honghan Wu

Noninvasive optical imaging modalities can probe patient's tissue in 3D and over time generate gigabytes of clinically relevant data per sample. There is a need for AI models to analyze this data and assist clinical workflow. The lack of…