English
Related papers

Related papers: A Multimodal Approach For Endoscopic VCE Image Cla…

200 papers

Joint image-text embedding extracted from medical images and associated contextual reports is the bedrock for most biomedical vision-and-language (V+L) tasks, including medical visual question answering, clinical image-text retrieval,…

Computer Vision and Pattern Recognition · Computer Science 2020-09-04 Yikuan Li , Hanyin Wang , Yuan Luo

The way we analyse clinical texts has undergone major changes over the last years. The introduction of language models such as BERT led to adaptations for the (bio)medical domain like PubMedBERT and ClinicalBERT. These models rely on large…

Computation and Language · Computer Science 2023-09-15 Tom van Sonsbeek , Xiantong Zhen , Marcel Worring

Wireless capsule endoscopy is a medical procedure used to visualize the entire gastrointestinal tract and to diagnose intestinal conditions, such as polyps or bleeding. Current analyses are performed by manually inspecting nearly each one…

Image and Video Processing · Electrical Eng. & Systems 2020-10-05 Pablo Laiz , Jordi Vitrià , Hagen Wenzek , Carolina Malagelada , Fernando Azpiroz , Santi Seguí

Accurate classification of focal liver lesions is crucial for diagnosis and treatment in hepatology. However, traditional supervised deep learning models depend on large-scale annotated datasets, which are often limited in medical imaging.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Song Jian , Hu Yuchang , Wang Hui , Chen Yen-Wei

We propose a novel pathology-sensitive deep learning model (PS-DeVCEM) for frame-level anomaly detection and multi-label classification of different colon diseases in video capsule endoscopy (VCE) data. Our proposed model is capable of…

Computer Vision and Pattern Recognition · Computer Science 2020-11-30 A. Mohammed , I. Farup , M. Pedersen , S. Yildirim , Ø Hovde

Capsule endoscopy (CE) enables non-invasive gastrointestinal screening, but current CE research remains largely limited to frame-level classification and detection, leaving video-level analysis underexplored. To bridge this gap, we…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Bowen Liu , Li Yang , Shanshan Song , Mingyu Tang , Zhifang Gao , Qifeng Chen , Yangqiu Song , Huimin Chen , Xiaomeng Li

In this paper, we construct two research objectives: i) explore the learned embedding space of BiomedCLIP, an open-source large vision language model, to analyse meaningful class separations, and ii) quantify the limitations of BiomedCLIP…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Nafiz Sadman , Farhana Zulkernine , Benjamin Kwan

Vision-language models and their adaptations to image segmentation tasks present enormous potential for producing highly accurate and interpretable results. However, implementations based on CLIP and BiomedCLIP are still lagging behind more…

Image and Video Processing · Electrical Eng. & Systems 2025-09-08 Julia Dietlmeier , Oluwabukola Grace Adegboro , Vayangi Ganepola , Claudia Mazo , Noel E. O'Connor

The escalating global mortality and morbidity rates associated with gastrointestinal (GI) bleeding, compounded by the complexities and limitations of traditional endoscopic methods, underscore the urgent need for a critical review of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Tanisha Singh , Shreshtha Jha , Nidhi Bhatt , Palak Handa , Nidhi Goel , Sreedevi Indu

CLIP and BiomedCLIP are examples of vision-language foundation models and offer strong cross-modal embeddings; however, they are not optimized for fine-grained medical retrieval tasks, such as retrieving clinically relevant radiology…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Zhaohui Liang , Sivaramakrishnan Rajaraman , Niccolo Marini , Zhiyun Xue , Sameer Antani

Accurate standard plane acquisition in fetal ultrasound (US) videos is crucial for fetal growth assessment, anomaly detection, and adherence to clinical guidelines. However, manually selecting standard frames is time-consuming and prone to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Divyanshu Mishra , Pramit Saha , He Zhao , Netzahualcoyotl Hernandez-Cruz , Olga Patey , Aris Papageorghiou , J. Alison Noble

Recent advancements in vision-language models (VLMs), such as CLIP, have demonstrated substantial success in self-supervised representation learning for vision tasks. However, effectively adapting VLMs to downstream applications remains…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Taha Koleilat , Hojat Asgariandehkordi , Hassan Rivaz , Yiming Xiao

Automated analysis of endoscopic imagery is a critical yet underdeveloped component of ENT (ear, nose, and throat) care, hindered by variability in devices and operators, subtle and localized findings, and fine-grained distinctions such as…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Trong-Thuan Nguyen , Viet-Tham Huynh , Thao Thi Phuong Dao , Ha Nguyen Thi , Tien To Vu Thuy , Uyen Hanh Tran , Tam V. Nguyen , Thanh Dinh Le , Minh-Triet Tran

Purpose. Early squamous cell neoplasia (ESCN) in the oesophagus is a highly treatable condition. Lesions confined to the mucosal layer can be curatively treated endoscopically. We build a computer-assisted detection (CADe) system that can…

Physicians use Capsule Endoscopy (CE) as a non-invasive and non-surgical procedure to examine the entire gastrointestinal (GI) tract for diseases and abnormalities. A single CE examination could last between 8 to 11 hours generating up to…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Sodiq Adewole , Philip Fernandes , James Jablonski , Andrew Copland , Michael Porter , Sana Syed , Donald Brown

In this paper we deal with image classification tasks using the powerful CLIP vision-language model. Our goal is to advance the classification performance using the CLIP's image encoder, by proposing a novel Large Multimodal Model (LMM)…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Maria Tzelepi , Vasileios Mezaris

Contrastive Language-Image Pretraining (CLIP) has demonstrated strong generalization for vision-language tasks in computer vision and medical domains, yet its text encoder accepts only up to 77 tokens, which limits its ability to represent…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Xiaoyang Wei , Camille Kurtz , Florence Cloppet

Video Capsule Endoscopy (VCE) is a promising method for improving the medical examination of the small intestine in the gastrointestinal tract. A key challenge is their limited size, resulting in a short battery lifetime which conflicts…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Oliver Bause , Jörg Gamerdinger , Julia Werner , Oliver Bringmann

In recent years, artificial intelligence has played an important role in medicine and disease diagnosis, with many applications to be mentioned, one of which is Medical Visual Question Answering (MedVQA). By combining computer vision and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Triet M. Thai , Anh T. Vo , Hao K. Tieu , Linh N. P. Bui , Thien T. B. Nguyen

Manual segmentation of medical images (e.g., segmenting tumors in CT scans) is a high-effort task that can be accelerated with machine learning techniques. However, selecting the right segmentation approach depends on the evaluation…

Computer Vision and Pattern Recognition · Computer Science 2023-02-09 Seyed M. R. Modaresi , Aomar Osmani , Mohammadreza Razzazi , Abdelghani Chibani