English
Related papers

Related papers: Assessing Multimodal Chronic Wound Embeddings with…

200 papers

Medical vision-language pretraining models (VLPM) have achieved remarkable progress in fusing chest X-rays (CXR) with clinical texts, introducing image-text data binding approaches that enable zero-shot learning and downstream clinical…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Yuan Gao , Sangwook Kim , David E Austin , Chris McIntosh

Diagnosing and treating skin diseases require advanced visual skills across domains and the ability to synthesize information from multiple imaging modalities. While current deep learning models excel at specific tasks like skin cancer…

This paper investigates the application of deep convolutional neural networks with prohibitively small datasets to the problem of macular edema segmentation. In particular, we investigate several different heavily regularized architectures.…

Image and Video Processing · Electrical Eng. & Systems 2020-05-12 Jonathan Frawley , Chris G. Willcocks , Maged Habib , Caspar Geenen , David H. Steel , Boguslaw Obara

Concept-based models naturally lend themselves to the development of inherently interpretable skin lesion diagnosis, as medical experts make decisions based on a set of visual patterns of the lesion. Nevertheless, the development of these…

Computer Vision and Pattern Recognition · Computer Science 2024-03-07 Cristiano Patrício , Luís F. Teixeira , João C. Neves

Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding robustness is multidimensional, since models respond differently to different types of…

Computation and Language · Computer Science 2026-05-28 Manuel Frank , Haithem Afli

Analyzing the morphology of cells in microscopy images can provide insights into the mechanism of compounds or the function of genes. Addressing this task requires methods that can not only extract biological information from the images,…

Machine Learning · Computer Science 2021-12-07 Siqi Wang , Manyuan Lu , Nikita Moshkov , Juan C. Caicedo , Bryan A. Plummer

Depression is a widespread mental health disorder, yet its automatic detection remains challenging. Prior work has explored unimodal and multimodal approaches, with multimodal systems showing promise by leveraging complementary signals.…

Artificial Intelligence · Computer Science 2026-03-24 Annisaa Fitri Nurfidausi , Eleonora Mancini , Paolo Torroni

Multimodal medical image fusion facilitates comprehensive diagnosis by aggregating complementary structural and functional information, but its effectiveness is limited by resolution degradation and modality discrepancies. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Fayaz Ali Dharejo , Sharif S. M. A. , Aiman Khalil , Nachiket Chaudhary , Rizwan Ali Naqvi , Radu Timofte

Many of the existing methods for learning joint embedding of images and text use only supervised information from paired images and its textual attributes. Taking advantage of the recent success of unsupervised learning in deep neural…

Computer Vision and Pattern Recognition · Computer Science 2017-03-21 Yao-Hung Hubert Tsai , Liang-Kang Huang , Ruslan Salakhutdinov

Diabetic retinopathy (DR) is a leading cause of preventable blindness worldwide, demanding accurate automated diagnostic systems. While general-domain vision-language models like Contrastive Language-Image Pre-Training (CLIP) perform well…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Argha Kamal Samanta , Harshika Goyal , Vasudha Joshi , Tushar Mungle , Pabitra Mitra

In the field of healthcare, precise skin lesion segmentation is crucial for the early detection and accurate diagnosis of skin diseases. Despite significant advances in deep learning for image processing, existing methods have yet to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Siyu Wang , Hua Wang , Huiyu Li , Fan Zhang

Out of all existing frameworks for surgical workflow analysis in endoscopic videos, action triplet recognition stands out as the only one aiming to provide truly fine-grained and comprehensive information on surgical activities. This…

Computer Vision and Pattern Recognition · Computer Science 2022-04-04 Chinedu Innocent Nwoye , Tong Yu , Cristians Gonzalez , Barbara Seeliger , Pietro Mascagni , Didier Mutter , Jacques Marescaux , Nicolas Padoy

Molecular property prediction aims to learn representations that map chemical structures to functional properties. While multimodal learning has emerged as a powerful paradigm to learn molecular representations, prior works have largely…

Machine Learning · Computer Science 2026-03-03 Feng Jiang , Mangal Prakash , Hehuan Ma , Jianyuan Deng , Yuzhi Guo , Amina Mollaysa , Tommaso Mansi , Rui Liao , Junzhou Huang

In this paper, the authors propose TriBERTa, a supervised entity resolution system that utilizes a pre-trained large language model and a triplet loss function to learn representations for entity matching. The system consists of two steps:…

Computation and Language · Computer Science 2024-11-19 Xiaowei Xu , Bi T. Foua , Xingqiao Wang , Vivek Gunasekaran , John R. Talburt

This paper discusses how ophthalmologists often rely on multimodal data to improve diagnostic accuracy. However, complete multimodal data is rare in real-world applications due to a lack of medical equipment and concerns about data privacy.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Xinkun Wang , Yifang Wang , Senwei Liang , Feilong Tang , Chengzhi Liu , Ming Hu , Chao Hu , Junjun He , Zongyuan Ge , Imran Razzak

Diffusion Probabilistic Models (DPMs) have demonstrated significant potential in 3D medical image segmentation tasks. However, their high computational cost and inability to fully capture global 3D contextual information limit their…

Image and Video Processing · Electrical Eng. & Systems 2025-04-17 Kangbo Ma

Skin lesion classification is essential for early dermatological diagnosis, yet many existing computer-aided systems rely primarily on dermoscopic images and underutilize the multimodal evidence routinely available in clinical practice. To…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Phan Nguyen , Dat Cao , Quang Hien Kha , Hien Chu , Minh H. N. Le , Trang Quoc Thao Pham , Nguyen Quoc Khanh Le

Deep learning and multi-modal fusion have demonstrated transformative potential in medical diagnosis by integrating diverse data sources. However, accurate prognosis for ischemic stroke remains challenging due to limitations in existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Liren Chen , Lidong Sun , Mingyan Huang , Junzhe Tang , Yinghui Zhu , Guanjie Wang , Yiqing Xia , Ting Xiao

Effective expression feature representations generated by a triplet-based deep metric learning are highly advantageous for facial expression recognition (FER). The performance of triplet-based deep metric learning is contingent upon…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Wenwu Yang , Jinyi Yu , Tuo Chen , Zhenguang Liu , Xun Wang , Jianbing Shen

The rapid advancement of Multimodal Large Language Models (MLLMs) has extended CLIP-based frameworks to produce powerful, universal embeddings for retrieval tasks. However, existing methods primarily focus on natural images, offering…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Weijian Jian , Yajun Zhang , Dawei Liang , Chunyu Xie , Yixiao He , Dawei Leng , Yuhui Yin
‹ Prev 1 2 3 10 Next ›