English
Related papers

Related papers: Swin SMT: Global Sequential Modeling in 3D Medical…

200 papers

This paper investigates the capability of plain Vision Transformers (ViTs) for semantic segmentation using the encoder-decoder framework and introduces \textbf{SegViTv2}. In this study, we introduce a novel Attention-to-Mask (\atm) module…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Bowen Zhang , Liyang Liu , Minh Hieu Phan , Zhi Tian , Chunhua Shen , Yifan Liu

Medical image segmentation plays an essential role in developing computer-assisted diagnosis and therapy systems, yet still faces many challenges. In the past few years, the popular encoder-decoder architectures based on CNNs (e.g., U-Net)…

Computer Vision and Pattern Recognition · Computer Science 2021-07-20 Guoping Xu , Xingrong Wu , Xuan Zhang , Xinwei He

In this paper, we present our approach to the Auto WCEBleedGen Challenge V2 2024. Our solution combines the Swin Transformer for the initial classification of bleeding frames and RT-DETR for further detection of bleeding in Wireless Capsule…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Sasidhar Alavala , Anil Kumar Vadde , Aparnamala Kancheti , Subrahmanyam Gorthi

Image analysis using more than one modality (i.e. multi-modal) has been increasingly applied in the field of biomedical imaging. One of the challenges in performing the multimodal analysis is that there exist multiple schemes for fusing the…

Computer Vision and Pattern Recognition · Computer Science 2018-06-19 Zhe Guo , Xiang Li , Heng Huang , Ning Guo , Quanzheng Li

Medical Image Segmentation (MIS) stands as a cornerstone in medical image analysis, playing a pivotal role in precise diagnostics, treatment planning, and monitoring of various medical conditions. This paper presents a comprehensive and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Ahmed Kabil , Ghada Khoriba , Mina Yousef , Essam A. Rashed

The conversion from 2D X-ray to 3D shape holds significant potential for improving diagnostic efficiency and safety. However, existing reconstruction methods often rely on hand-crafted features, manual intervention, and prior knowledge,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Kuan Liu , Zongyuan Ying , Jie Jin , Dongyan Li , Ping Huang , Wenjian Wu , Zhe Chen , Jin Qi , Yong Lu , Lianfu Deng , Bo Chen

Providing more precise tissue attenuation information, synthetic computed tomography (sCT) generated from magnetic resonance imaging (MRI) contributes to improved radiation therapy treatment planning. In our study, we employ the advanced…

Image and Video Processing · Electrical Eng. & Systems 2024-09-11 Fuxin Fan , Jingna Qiu , Yixing Huang , Andreas Maier

Pansharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multispectral (LRMS) images to generate high-resolution multispectral (HRMS) images. Although deep learning-based methods have achieved promising…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zeyu Xia , Chenxi Sun , Tianyu Xin , Yubo Zeng , Haoyu Chen , Liang-Jian Deng

In the paper, we present an approach for learning a single model that universally segments 33 anatomical structures, including vertebrae, pelvic bones, and abdominal organs. Our model building has to address the following challenges.…

Image and Video Processing · Electrical Eng. & Systems 2022-03-07 Pengbo Liu , Yang Deng , Ce Wang , Yuan Hui , Qian Li , Jun Li , Shiwei Luo , Mengke Sun , Quan Quan , Shuxin Yang , You Hao , Honghu Xiao , Chunpeng Zhao , Xinbao Wu , S. Kevin Zhou

Vision Transformers (ViTs) have shown promise in medical image semantic segmentation (MISS) by capturing long-range correlations. However, ViTs often struggle to model local spatial information effectively, which is essential for accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Niloufar Eghbali , Hassan Bagher-Ebadian , Tuka Alhanai , Mohammad M. Ghassemi

In this paper, we propose a weakly supervised semantic segmentation approach for food images which takes advantage of the zero-shot capabilities and promptability of the Segment Anything Model (SAM) along with the attention mechanisms of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Ioannis Sarafis , Alexandros Papadopoulos , Anastasios Delopoulos

Simulation studies, such as finite element (FE) modeling, offer insights into knee joint biomechanics, which may not be achieved through experimental methods without direct involvement of patients. While generic FE models have been used to…

Image and Video Processing · Electrical Eng. & Systems 2024-07-10 Reza Kakavand , Peyman Tahghighi , Reza Ahmadi , W. Brent Edwards , Amin Komeili

Self-supervised learning (SSL) opens up huge opportunities for medical image analysis that is well known for its lack of annotations. However, aggregating massive (unlabeled) 3D medical images like computerized tomography (CT) remains…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Yutong Xie , Jianpeng Zhang , Yong Xia , Qi Wu

Rapid advancements in medical image segmentation performance have been significantly driven by the development of Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs). These models follow the discriminative pixel-wise…

Image and Video Processing · Electrical Eng. & Systems 2024-08-21 Jiayu Huo , Xi Ouyang , Sébastien Ourselin , Rachel Sparks

Tissue window filtering has been widely used in deep learning for computed tomography (CT) image analyses to improve training performance (e.g., soft tissue windows for abdominal CT). However, the effectiveness of tissue window…

Image and Video Processing · Electrical Eng. & Systems 2019-12-03 Yuankai Huo , Yucheng Tang , Yunqiang Chen , Dashan Gao , Shizhong Han , Shunxing Bao , Smita De , James G. Terry , Jeffrey J. Carr , Richard G. Abramson , Bennett A. Landman

Transformer models have shown great potential in computer vision, following their success in language tasks. Swin Transformer is one of them that outperforms convolution-based architectures in terms of accuracy, while improving efficiency…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Jinkyu Koo , John Yang , Le An , Gwenaelle Cunha Sergio , Su Inn Park

Convolutional Neural Networks (CNNs) for computer vision sometimes struggle with understanding images in a global context, as they mainly focus on local patterns. On the other hand, Vision Transformers (ViTs), inspired by models originally…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Dimitrios N. Vlachogiannis , Dimitrios A. Koutsomitropoulos

In recent years, Vision Transformer-based approaches for low-level vision tasks have achieved widespread success. Unlike CNN-based models, Transformers are more adept at capturing long-range dependencies, enabling the reconstruction of…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Chih-Chung Hsu , Chia-Ming Lee , Yi-Shiuan Chou

Skin cancer segmentation poses a significant challenge in medical image analysis. Numerous existing solutions, predominantly CNN-based, face issues related to a lack of global contextual understanding. Alternatively, some approaches resort…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Shehan Perera , Yunus Erzurumlu , Deepak Gulati , Alper Yilmaz

In the field of medical image segmentation, variant models based on Convolutional Neural Networks (CNNs) and Visual Transformers (ViTs) as the base modules have been very widely developed and applied. However, CNNs are often limited in…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Renkai Wu , Yinghao Liu , Pengchen Liang , Qing Chang