English
Related papers

Related papers: ConTEXTual Net: A Multimodal Vision-Language Model…

200 papers

Vision-Language Pre-training (VLP) is drawing increasing interest for its ability to minimize manual annotation requirements while enhancing semantic understanding in downstream tasks. However, its reliance on image-text datasets poses…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Sinuo Wang , Yutong Xie , Yuyuan Liu , Qi Wu

It is a challenge to segment the location and size of rectal cancer tumours through deep learning. In this paper, in order to improve the ability of extracting suffi-cient feature information in rectal tumour segmentation, attention…

Image and Video Processing · Electrical Eng. & Systems 2022-10-28 Hongwei Wu , Junlin Wang , Xin Wang , Hui Nan , Yaxin Wang , Haonan Jing , Kaixuan Shi

We propose a unified optimization framework that combines neural networks with dictionary learning to model complex interactions between resting state functional MRI and behavioral data. The dictionary learning objective decomposes patient…

Machine Learning · Computer Science 2024-11-21 Niharika Shimona D'Souza , Mary Beth Nebel , Nicholas Wymbs , Stewart Mostofsky , Archana Venkataraman

Deep learning has shown great potential for automated medical image segmentation to improve the precision and speed of disease diagnostics. However, the task presents significant difficulties due to variations in the scale, shape, texture,…

Image and Video Processing · Electrical Eng. & Systems 2024-09-06 Shahzaib Iqbal , Tariq M. Khan , Syed S. Naqvi , Asim Naveed , Erik Meijering

The utilisation of deep learning segmentation algorithms that learn complex organs and tissue patterns and extract essential regions of interest from the noisy background to improve the visual ability for medical image diagnosis has…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Yanming Guo

Background: Precise breast ultrasound (BUS) segmentation supports reliable measurement, quantitative analysis, and downstream classification, yet remains difficult for small or low-contrast lesions with fuzzy margins and speckle noise. Text…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Raja Mallina , Bryar Shareef

The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks…

Accurate retinal vessel segmentation is a challenging problem in color fundus image analysis. An automatic retinal vessel segmentation system can effectively facilitate clinical diagnosis and ophthalmological research. Technically, this…

Image and Video Processing · Electrical Eng. & Systems 2021-03-26 Muyi Sun , Guanhong Zhang

Detecting clinically relevant objects in medical images is a challenge despite large datasets due to the lack of detailed labels. To address the label issue, we utilize the scene-level labels with a detection architecture that incorporates…

Computer Vision and Pattern Recognition · Computer Science 2020-08-03 Leo K. Tam , Xiaosong Wang , Evrim Turkbey , Kevin Lu , Yuhong Wen , Daguang Xu

Recently, vision-language joint representation learning has proven to be highly effective in various scenarios. In this paper, we specifically adapt vision-language joint learning for scene text detection, a task that intrinsically involves…

Computer Vision and Pattern Recognition · Computer Science 2022-05-02 Sibo Song , Jianqiang Wan , Zhibo Yang , Jun Tang , Wenqing Cheng , Xiang Bai , Cong Yao

Thoracic disease detection from chest radiographs using deep learning methods has been an active area of research in the last decade. Most previous methods attempt to focus on the diseased organs of the image by identifying spatial regions…

Image and Video Processing · Electrical Eng. & Systems 2022-10-07 Uday Kamal , Mohammad Zunaed , Nusrat Binta Nizam , Taufiq Hasan

Most current medical vision language models struggle to jointly generate diagnostic text and pixel-level segmentation masks in response to complex visual questions. This represents a major limitation towards clinical application, as…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Chengrun Li , Corentin Royer , Haozhe Luo , Bastian Wittmann , Xia Li , Ibrahim Hamamci , Sezgin Er , Anjany Sekuboyina , Bjoern Menze

Diagnosing dental diseases from radiographs is time-consuming and challenging due to the subtle nature of diagnostic evidence. Existing methods, which rely on object detection models designed for natural images with more distinct target…

Computer Vision and Pattern Recognition · Computer Science 2026-01-14 Zhi Qin Tan , Xiatian Zhu , Owen Addison , Yunpeng Li

Pneumothorax, the abnormal accumulation of air in the pleural space, can be life-threatening if undetected. Chest X-rays are the first-line diagnostic tool, but small cases may be subtle. We propose an automated deep-learning pipeline using…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Alvaro Aranibar Roque , Helga Sebastian

Pneumonia remains a leading global cause of mortality where timely diagnosis is critical. We introduce LungX, a novel hybrid architecture combining EfficientNet's multi-scale features, CBAM attention mechanisms, and Vision Transformer's…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Mansur Yerzhanuly

Volumetric medical segmentation is a critical component of 3D medical image analysis that delineates different semantic regions. Deep neural networks have significantly improved volumetric medical segmentation, but they generally require…

Image and Video Processing · Electrical Eng. & Systems 2024-07-18 Hanan Gani , Muzammal Naseer , Fahad Khan , Salman Khan

The precise segmentation of retinal blood vessels is of great significance for early diagnosis of eye-related diseases such as diabetes and hypertension. In this work, we propose a lightweight network named Spatial Attention U-Net (SA-UNet)…

Image and Video Processing · Electrical Eng. & Systems 2020-10-22 Changlu Guo , Márton Szemenyei , Yugen Yi , Wenle Wang , Buer Chen , Changqi Fan

We introduce VoxTell, a vision-language model for text-prompted volumetric medical image segmentation. It maps free-form descriptions, from single words to full clinical sentences, to 3D masks. Trained on 62K+ CT, MRI, and PET volumes…

Purpose: To investigate chest radiograph (CXR) classification performance of vision transformers (ViT) and interpretability of attention-based saliency using the example of pneumothorax classification. Materials and Methods: In this…

Image and Video Processing · Electrical Eng. & Systems 2023-05-09 Alessandro Wollek , Robert Graf , Saša Čečatka , Nicola Fink , Theresa Willem , Bastian O. Sabel , Tobias Lasser

Semantic segmentation in surgical videos is a prerequisite for a broad range of applications towards improving surgical outcomes and surgical video analysis. However, semantic segmentation in surgical videos involves many challenges. In…

Image and Video Processing · Electrical Eng. & Systems 2021-09-28 Negin Ghamsarian , Mario Taschwer , Doris Putzgruber-Adamitsch , Stephanie Sarny , Yosuf El-Shabrawi , Klaus Schoeffmann