English
Related papers

Related papers: Self-supervised vision-langage alignment of deep l…

200 papers

Deep learning in medical imaging has the potential to minimize the risk of diagnostic errors, reduce radiologist workload, and accelerate diagnosis. Training such deep learning models requires large and accurate datasets, with annotations…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Daniel Wolf , Tristan Payer , Catharina Silvia Lisson , Christoph Gerhard Lisson , Meinrad Beer , Michael Götz , Timo Ropinski

Writing radiology reports from medical images requires a high level of domain expertise. It is time-consuming even for trained radiologists and can be error-prone for inexperienced radiologists. It would be appealing to automate this task…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Yuzhe Lu , Sungmin Hong , Yash Shah , Panpan Xu

Medical vision-language pretraining increasingly relies on medical reports as large-scale supervisory signals; however, raw reports often exhibit substantial stylistic heterogeneity, variable length, and a considerable amount of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Yuetan Chu , Xinhua Ma , Xinran Jin , Gongning Luo , Xin Gao

Self-supervised pre-training appears as an advantageous alternative to supervised pre-trained for transfer learning. By synthesizing annotations on pretext tasks, self-supervision allows to pre-train models on large amounts of pseudo-labels…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Levy Chaves , Alceu Bissoto , Eduardo Valle , Sandra Avila

We present a model that generates natural language descriptions of images and their regions. Our approach leverages datasets of images and their sentence descriptions to learn about the inter-modal correspondences between language and…

Computer Vision and Pattern Recognition · Computer Science 2015-04-15 Andrej Karpathy , Li Fei-Fei

Medical vision-language pre-training shows great potential in learning representative features from massive paired radiographs and reports. However, in computed tomography (CT) scans, the distribution of lesions which contain intricate…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Rongsheng Wang , Fenghe Tang , Qingsong Yao , Rui Yan , Xu Zhang , Zhen Huang , Haoran Lai , Zhiyang He , Xiaodong Tao , Zihang Jiang , Shaohua Kevin Zhou

With the development of multimodality and large language models, the deep learning-based technique for medical image captioning holds the potential to offer valuable diagnostic recommendations. However, current generic text and image…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Zhenyu Zhang , Benlu Wang , Weijie Liang , Yizhi Li , Xuechen Guo , Guanhong Wang , Shiyan Li , Gaoang Wang

Recently, self-supervised learning (SSL) methods have been used in pre-training the segmentation models for 2D and 3D medical images. Most of these methods are based on reconstruction, contrastive learning and consistency regularization.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Haofeng Li , Yiming Ouyang , Xiang Wan

Computer-aided systems in histopathology are often challenged by various sources of domain shift that impact the performance of these algorithms considerably. We investigated the potential of using self-supervised pre-training to overcome…

Recent advances in vision-language models (VLMs) have shown remarkable potential in bridging visual and textual modalities. In computational pathology, domain-specific VLMs, which are pre-trained on extensive histopathology image-text…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Anh Tien Nguyen , Keunho Byeon , Kyungeun Kim , Jin Tae Kwak

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-21 Abhinav Shukla , Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

Integrating information from multiple modalities is arguably one of the essential prerequisites for grounding artificial intelligence systems with an understanding of the real world. Recent advances in video transformers that jointly learn…

Computer Vision and Pattern Recognition · Computer Science 2023-11-15 Dota Tianai Dong , Mariya Toneva

Background: Existing clinical prediction models often represent patient data using features that ignore the semantic relationships between clinical concepts. This study integrates domain-specific semantic information by mapping the SNOMED…

Machine Learning · Computer Science 2025-08-21 Luis H. John , Jan A. Kors , Jenna M. Reps , Peter R. Rijnbeek , Egill A. Fridgeirsson

Vision-language models in pathology enable multimodal case retrieval and automated report generation. Many of the models developed so far, however, have been trained on pathology reports that include information which cannot be inferred…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Ruben T. Lucassen , Tijn van de Luijtgaarden , Sander P. J. Moonemans , Gerben E. Breimer , Willeke A. M. Blokx , Mitko Veta

Objective and Impact Statement. We adopt a deep learning model for bone osteolysis prediction on computed tomography (CT) images of murine breast cancer bone metastases. Given the bone CT scans at previous time steps, the model incorporates…

Image and Video Processing · Electrical Eng. & Systems 2022-03-29 Wei Xiong , Neil Yeung , Shubo Wang , Haofu Liao , Liyun Wang , Jiebo Luo

The diagnosis of primary bone tumors is challenging, as the initial complaints are often non-specific. Early detection of bone cancer is crucial for a favorable prognosis. Incidentally, lesions may be found on radiographs obtained for other…

Image and Video Processing · Electrical Eng. & Systems 2024-03-22 Tal Zimbalist , Ronnie Rosen , Keren Peri-Hanania , Yaron Caspi , Bar Rinott , Carmel Zeltser-Dekel , Eyal Bercovich , Yonina C. Eldar , Shai Bagon

Deep learning technologies have already demonstrated a high potential to build diagnosis support systems from medical imaging data, such as Chest X-Ray images. However, the shortage of labeled data in the medical field represents one key…

Image and Video Processing · Electrical Eng. & Systems 2023-01-26 Iván de Andrés Tamé , Kirill Sirotkin , Pablo Carballeira , Marcos Escudero-Viñolo

An important challenge in texture recognition is the limited amount of data for training frequently found in real-world applications. In computer vision in general, a successful strategy to mitigate this issue is the use of a pretraining…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Joao B. Florindo , Lucas O. Lyra , Antonio E. Fabris

Localization and characterization of diseases like pneumonia are primary steps in a clinical pipeline, facilitating detailed clinical diagnosis and subsequent treatment planning. Additionally, such location annotated datasets can provide a…

Image and Video Processing · Electrical Eng. & Systems 2021-10-08 Riddhish Bhalodia , Ali Hatamizadeh , Leo Tam , Ziyue Xu , Xiaosong Wang , Evrim Turkbey , Daguang Xu

Transfer learning has become a standard practice to mitigate the lack of labeled data in medical classification tasks. Whereas finetuning a downstream task using supervised ImageNet pretrained features is straightforward and extensively…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Tuan Truong , Sadegh Mohammadi , Matthias Lenga
‹ Prev 1 8 9 10 Next ›