中文
相关论文

相关论文: CapCLIP: A Vision-Language Representation Alignmen…

200 篇论文

We propose a novel pathology-sensitive deep learning model (PS-DeVCEM) for frame-level anomaly detection and multi-label classification of different colon diseases in video capsule endoscopy (VCE) data. Our proposed model is capable of…

计算机视觉与模式识别 · 计算机科学 2020-11-30 A. Mohammed , I. Farup , M. Pedersen , S. Yildirim , Ø Hovde

Video Capsule Endoscopy (VCE) is a promising method for improving the medical examination of the small intestine in the gastrointestinal tract. A key challenge is their limited size, resulting in a short battery lifetime which conflicts…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Oliver Bause , Jörg Gamerdinger , Julia Werner , Oliver Bringmann

Large-scale vision-language models such as CLIP have achieved remarkable success in zero-shot image recognition, yet their predictions remain largely opaque to human understanding. In contrast, Concept Bottleneck Models provide…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Onat Ozdemir , Anders Christensen , Stephan Alaniz , Zeynep Akata , Emre Akbas

Capsule endoscopes, predominantly serving diagnostic functions, provide lucid internal imagery but are devoid of surgical or therapeutic capabilities. Consequently, despite lesion detection, physicians frequently resort to traditional…

医学物理 · 物理学 2024-05-14 Sishen Yuan , Guang Li , Baijia Liang , Lailu Li , Qingzhuo Zheng , Shuang Song , Zhen Li , Hongliang Ren

We present a unified vision-language framework tailored for ENT endoscopy image analysis that simultaneously tackles three clinically-relevant tasks: image classification, image-to-image retrieval, and text-to-image retrieval. Unlike…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Y Hop Nguyen , Doan Anh Phan Huu , Trung Thai Tran , Nhat Nam Mai , Van Toi Giap , Thao Thi Phuong Dao , Trung-Nghia Le

Temporal activity localization in long videos is an important problem. The cost of obtaining frame level label for long Wireless Capsule Endoscopy (WCE) videos is prohibitive. In this paper, we propose an end-to-end temporal abnormality…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Sodiq Adewole , Philip Fernandes , James Jablonski , Andrew Copland , Michael Porter , Sana Syed , Donald Brown

General-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation models. However, these…

计算机视觉与模式识别 · 计算机科学 2024-04-17 Fan Liu , Delong Chen , Zhangqingyun Guan , Xiaocong Zhou , Jiale Zhu , Qiaolin Ye , Liyong Fu , Jun Zhou

The integration of artificial intelligence (AI) with radiology marks a transformative era in medicine. Vision foundation models have been adopted to enhance radiologic imaging analysis. However, the distinct complexities of radiologic 2D…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Zhixiu Lu , Hailong Li , Nehal A. Parikh , Jonathan R. Dillman , Lili He

Dense visual perception tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Junjie Wang , Keyu Chen , Yulin Li , Bin Chen , Hengshuang Zhao , Xiaojuan Qi , Zhuotao Tian

We introduce WAVECLIP, a single unified model for adaptive resolution inference in CLIP, enabled by wavelet-based tokenization. WAVECLIP replaces standard patch embeddings with a multi-level wavelet decomposition, enabling the model to…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Moshe Kimhi , Erez Koifman , Ehud Rivlin , Eli Schwartz , Chaim Baskin

Wireless Capsule Endoscopy (WCE) is highly valued for its non-invasive and painless approach, though its effectiveness is compromised by uneven illumination from hardware constraints and complex internal dynamics, leading to overexposed or…

图像与视频处理 · 电气工程与系统科学 2024-07-09 Long Bai , Tong Chen , Qiaozhi Tan , Wan Jun Nah , Yanheng Li , Zhicheng He , Sishen Yuan , Zhen Chen , Jinlin Wu , Mobarakol Islam , Zhen Li , Hongbin Liu , Hongliang Ren

Wireless capsule endoscopy (WCE) enables painless visualization of the gastrointestinal tract, but its diagnostic potential is limited by incomplete mucosal coverage and poor transferability of existing navigation methods across patient…

机器人学 · 计算机科学 2026-05-12 Haoxuan Wu , Sishen Yuan , Haitao Gao , Zhen Li , Xiuli Zuo , Hongliang Ren

In diagnostic reports, experts encode complex imaging data into clinically actionable information. They describe subtle pathological findings that are meaningful in their anatomical context. Reports follow relatively consistent structures,…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Felicia Bader , Philipp Seeböck , Anastasia Bartashova , Ulrike Attenberger , Georg Langs

CLIP embeddings have demonstrated remarkable performance across a wide range of multimodal applications. However, these high-dimensional, dense vector representations are not easily interpretable, limiting our understanding of the rich…

机器学习 · 计算机科学 2024-11-05 Usha Bhalla , Alex Oesterling , Suraj Srinivas , Flavio P. Calmon , Himabindu Lakkaraju

In the last decade, many medical companies and research groups have tried to convert passive capsule endoscopes as an emerging and minimally invasive diagnostic technology into actively steerable endoscopic capsule robots which will provide…

Vision-language models such as CLIP are capable of mapping the different modality data into a unified feature space, enabling zero/few-shot inference by measuring the similarity of given images and texts. However, most existing methods…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Xingyu Zhu , Beier Zhu , Yi Tan , Shuo Wang , Yanbin Hao , Hanwang Zhang

Physicians use Capsule Endoscopy (CE) as a non-invasive and non-surgical procedure to examine the entire gastrointestinal (GI) tract for diseases and abnormalities. A single CE examination could last between 8 to 11 hours generating up to…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Sodiq Adewole , Philip Fernandes , James Jablonski , Andrew Copland , Michael Porter , Sana Syed , Donald Brown

We present CapsoNet, a deep learning framework developed for the Capsule Vision 2024 Challenge, designed to perform multi-class abnormality classification in video capsule endoscopy (VCE) frames. CapsoNet leverages an ensemble of…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Arnav Samal , Ranya Batsyas

Large vision-language contrastive models (VLCMs), such as CLIP, have become foundational, demonstrating remarkable success across a variety of downstream tasks. Despite their advantages, these models, akin to other foundational systems,…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Haocheng Dai , Sarang Joshi

This paper presents, for the first time, an image enhancement methodology designed to enhance the clarity of small intestinal villi in Wireless Capsule Endoscopy (WCE) images. This method first separates the low-frequency and high-frequency…

计算机视觉与模式识别 · 计算机科学 2024-10-25 Shaojie Zhang , Yinghui Wang , Peixuan Liu , Yukai Wang , Liangyi Huang , Mingfeng Wang , Ibragim. Atadjanov