English
Related papers

Related papers: ZACH-ViT: A Zero-Token Vision Transformer with Shu…

200 papers

In this study, we introduce an intelligent Test Time Augmentation (TTA) algorithm designed to enhance the robustness and accuracy of image classification models against viewpoint variations. Unlike traditional TTA methods that…

Image and Video Processing · Electrical Eng. & Systems 2024-06-14 Efe Ozturk , Mohit Prabhushankar , Ghassan AlRegib

Optical Coherence Tomography (OCT) provides high-resolution cross-sectional images useful for diagnosing various diseases, but their distinct characteristics from natural images raise questions about whether large-scale pre-training on…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Zihao Han , Philippe De Wilde

Manual labeling of animal images remains a significant bottleneck in ecological research, limiting the scale and efficiency of biodiversity monitoring efforts. This study investigates whether state-of-the-art Vision Transformer (ViT)…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Hugo Markoff , Stefan Hein Bengtson , Michael Ørsted

Video anomaly detection (VAD) aims to identify abnormal events in videos. Traditional VAD methods generally suffer from the high costs of labeled data and full training, thus some recent works have explored leveraging frozen multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Zhaolin Cai , Fan Li , Huiyu Duan , Lijun He , Guangtao Zhai

Medical vision-language alignment through cross-modal contrastive learning shows promising performance in image-text matching tasks, such as retrieval and zero-shot classification. However, conventional cross-modal contrastive learning…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Chenyu Lian , Hong-Yu Zhou , Dongyun Liang , Jing Qin , Liansheng Wang

Although Video Large Language Models (VLLMs) have shown remarkable capabilities in video understanding, they are required to process high volumes of visual tokens, causing significant computational inefficiency. Existing VLLMs acceleration…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Ziyang Fan , Keyu Chen , Ruilong Xing , Yulin Li , Li Jiang , Zhuotao Tian

Vision Transformer (ViT) has shown great potential for various visual tasks due to its ability to model long-range dependency. However, ViT requires a large amount of computing resource to compute the global self-attention. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Gaojie Wu , Wei-Shi Zheng , Yutong Lu , Qi Tian

Accurate segmentation of cardiac substructures on computed tomography (CT) scans is essential for radiotherapy planning but typically requires large annotated datasets and often generalizes poorly across imaging protocols and patient…

Image and Video Processing · Electrical Eng. & Systems 2026-02-26 Aneesh Rangnekar , Nikhil Mankuzhy , Jonas Willmann , Chloe Min Seo Choi , Abraham Wu , Maria Thor , Andreas Rimner , Harini Veeraraghavan

In the last decade, convolutional neural networks (ConvNets) have dominated and achieved state-of-the-art performances in a variety of medical imaging applications. However, the performances of ConvNets are still limited by lacking the…

Image and Video Processing · Electrical Eng. & Systems 2021-04-15 Junyu Chen , Yufan He , Eric C. Frey , Ye Li , Yong Du

Robust object recognition systems usually rely on powerful feature extraction mechanisms from a large number of real images. However, in many realistic applications, collecting sufficient images for ever-growing new classes is unattainable.…

Computer Vision and Pattern Recognition · Computer Science 2017-05-05 Yang Long , Li Liu , Ling Shao , Fumin Shen , Guiguang Ding , Jungong Han

In recent years, deep learning has made brilliant achievements in Environmental Microorganism (EM) image classification. However, image classification of small EM datasets has still not obtained good research results. Therefore, researchers…

Computer Vision and Pattern Recognition · Computer Science 2022-02-04 Peng Zhao , Chen Li , Md Mamunur Rahaman , Hao Xu , Hechen Yang , Hongzan Sun , Tao Jiang , Marcin Grzegorzek

Medical image analysis is a hot research topic because of its usefulness in different clinical applications, such as early disease diagnosis and treatment. Convolutional neural networks (CNNs) have become the de-facto standard in medical…

Image and Video Processing · Electrical Eng. & Systems 2023-04-25 Smriti Regmi , Aliza Subedi , Ulas Bagci , Debesh Jha

The interconnection between the human lungs and other organs, such as the liver and kidneys, is crucial for understanding the underlying risks and effects of lung diseases and improving patient care. However, most research chest CT imaging…

Computer Vision and Pattern Recognition · Computer Science 2025-01-23 Lianrui Zuo , Kaiwen Xu , Dingjie Su , Xin Yu , Aravind R. Krishnan , Yihao Liu , Shunxing Bao , Thomas Li , Kim L. Sandler , Fabien Maldonado , Bennett A. Landman

Vision transformer (ViT) has been widely applied in many areas due to its self-attention mechanism that help obtain the global receptive field since the first layer. It even achieves surprising performance exceeding CNN in some vision…

Computer Vision and Pattern Recognition · Computer Science 2021-09-28 Hanting Li , Mingzhe Sui , Zhaoqing Zhu , Feng Zhao

Localization of the narrowest position of the vessel and corresponding vessel and remnant vessel delineation in carotid ultrasound (US) are essential for carotid stenosis grading (CSG) in clinical practice. However, the pipeline is…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Xinrui Zhou , Yuhao Huang , Wufeng Xue , Xin Yang , Yuxin Zou , Qilong Ying , Yuanji Zhang , Jia Liu , Jie Ren , Dong Ni

Vision Transformers (ViTs) have achieved remarkable success in standard RGB image processing tasks. However, applying ViTs to multi-channel imaging (MCI) data, e.g., for medical and remote sensing applications, remains a challenge. In…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Wenyi Lian , Patrick Micke , Joakim Lindblad , Nataša Sladoje

Vision Transformer (ViT) attains state-of-the-art performance in visual recognition, and the variant, Local Vision Transformer, makes further improvements. The major component in Local Vision Transformer, local attention, performs the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-05 Qi Han , Zejia Fan , Qi Dai , Lei Sun , Ming-Ming Cheng , Jiaying Liu , Jingdong Wang

Accurate identification of late-life depression (LLD) using structural brain MRI is essential for monitoring disease progression and facilitating timely intervention. However, existing learning-based approaches for LLD detection are often…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Yuzhen Gao , Qianqian Wang , Yongheng Sun , Cui Wang , Yongquan Liang , Mingxia Liu

Advanced diagnostic instruments are crucial for the accurate detection and treatment of lung diseases, which affect millions of individuals globally. This study examines the effectiveness of deep learning and transfer learning models using…

Image and Video Processing · Electrical Eng. & Systems 2025-06-23 Shuvashis Sarker , Shamim Rahim Refat , Faika Fairuj Preotee , Tanvir Rouf Shawon , Raihan Tanvir

Abnormalities in the gastrointestinal tract significantly influence the patient's health and require a timely diagnosis for effective treatment. With such consideration, an effective automatic classification of these abnormalities from a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Lakshmi Srinivas Panchananam , Praveen Kumar Chandaliya , Kishor Upla , Kiran Raja
‹ Prev 1 8 9 10 Next ›