English
Related papers

Related papers: Dual-Domain CLIP-Assisted Residual Optimization Pe…

200 papers

We propose DiffCLIP, a novel vision-language model that extends the differential attention mechanism to CLIP architectures. Differential attention was originally developed for large language models to amplify relevant context while…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Hasan Abed Al Kader Hammoud , Bernard Ghanem

Zero-shot 3D object classification is crucial for real-world applications like autonomous driving, however it is often hindered by a significant domain gap between the synthetic data used for training and the sparse, noisy LiDAR scans…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Ajinkya Khoche , Gergő László Nagy , Maciej Wozniak , Thomas Gustafsson , Patric Jensfelt

We propose approaches based on deep learning to localize objects in images when only a small training dataset is available and the images have low quality. That applies to many problems in medical image processing, and in particular to the…

Computer Vision and Pattern Recognition · Computer Science 2018-11-20 Aaron Pries , Peter J. Schreier , Artur Lamm , Stefan Pede , Jürgen Schmidt

Undersampled MRI reconstruction is crucial for accelerating clinical scanning. Dual-domain reconstruction network is performant among SoTA deep learning methods. In this paper, we rethink dual-domain model design from the perspective of the…

Image and Video Processing · Electrical Eng. & Systems 2024-02-16 Ziqi Gao , S. Kevin Zhou

Few-shot Generalist Anomaly Detection requires models to generalize to novel categories without retraining, posing significant challenges in real-world scenarios with scarce samples and rapidly changing categories. Existing CLIP-based…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Xinyue Liu , Jianyuan Wang , Biao Leng , Shuo Zhang

Metal artifacts is a major challenge in computed tomography (CT) imaging, significantly degrading image quality and making accurate diagnosis difficult. However, previous methods either require prior knowledge of the location of metal…

Image and Video Processing · Electrical Eng. & Systems 2023-06-21 Jiandong Su , Ce Wang , Yinsheng Li , Kun Shang , Dong Liang

Semi-supervised object detection methods are widely used in autonomous driving systems, where only a fraction of objects are labeled. To propagate information from the labeled objects to the unlabeled ones, pseudo-labels for unlabeled…

Computer Vision and Pattern Recognition · Computer Science 2022-04-19 Shu Hu , Chun-Hao Liu , Jayanta Dutta , Ming-Ching Chang , Siwei Lyu , Naveen Ramakrishnan

In computed tomography imaging, metal implants frequently generate severe artifacts that compromise image quality and hinder diagnostic accuracy. There are three main challenges in the existing methods: the deterioration of organ and tissue…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Weikai Qu , Sijun Liang , Xianfeng Li , Cheng Pan , An Yan , Ahmed Elazab , Shanzhou Niu , Dong Zeng , Xiang Wan , Changmiao Wang

In medical imaging, most of the image registration methods implicitly assume a one-to-one correspondence between the source and target images (i.e., diffeomorphism). However, this is not necessarily the case when dealing with pathological…

Image and Video Processing · Electrical Eng. & Systems 2022-02-03 Matthis Maillard , Anton François , Joan Glaunès , Isabelle Bloch , Pietro Gori

Foundation models have recently gained tremendous popularity in medical image analysis. State-of-the-art methods leverage either paired image-text data via vision-language pre-training or unpaired image data via self-supervised pre-training…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Lei Zhu , Jun Zhou , Rick Siow Mong Goh , Yong Liu

In total hip arthroplasty, analysis of postoperative medical images is important to evaluate surgical outcome. Since Computed Tomography (CT) is most prevalent modality in orthopedic surgery, we aimed at the analysis of CT image. In this…

Image and Video Processing · Electrical Eng. & Systems 2019-06-28 Mitsuki Sakamoto , Yuta Hiasa , Yoshito Otake , Masaki Takao , Yuki Suzuki , Nobuhiko Sugano , Yoshinobu Sato

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels typically leverages Class Activation Maps (CAMs) to achieve pixel-level predictions. Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Zhiwei Yang , Pengfei Song , Yucong Meng , Kexue Fu , Shuo Wang , Zhijian Song

Diffusion-weighted MRI is nowadays performed routinely due to its prognostic ability, yet the quality of the scans are often unsatisfactory which can subsequently hamper the clinical utility. To overcome the limitations, here we propose a…

Image and Video Processing · Electrical Eng. & Systems 2021-05-04 Hyungjin Chung , Jaehyun Kim , Jeong Hee Yoon , Jeong Min Lee , Jong Chul Ye

While multi-modal Visual Language Models (VLMs) have demonstrated significant success across various domains, the integration of VLMs into recommendation and retrieval systems remains a challenge, due to issues like training objective…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Josh Beal , Eric Kim , Jinfeng Rao , Rex Wu , Dmitry Kislyuk , Charles Rosenberg

Signal models based on sparse representations have received considerable attention in recent years. On the other hand, deep models consisting of a cascade of functional layers, commonly known as deep neural networks, have been highly…

Image and Video Processing · Electrical Eng. & Systems 2022-01-19 Xikai Yang , Yong Long , Saiprasad Ravishankar

Image colorization is a challenging problem due to multi-modal uncertainty and high ill-posedness. Directly training a deep neural network usually leads to incorrect semantic colors and low color richness. While transformer-based methods…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Xiaoyang Kang , Tao Yang , Wenqi Ouyang , Peiran Ren , Lingzhi Li , Xuansong Xie

In the field of design patent analysis, traditional tasks such as patent classification and patent image retrieval heavily depend on the image data. However, patent images -- typically consisting of sketches with abstract and structural…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Zhu Wang , Homaira Huda Shomee , Sathya N. Ravi , Sourav Medya

High performance materials, from natural bone over ancient damascene steel to modern superalloys, typically possess a complex structure at the microscale. Their properties exceed those of the individual components and their knowledge-based…

Materials Science · Physics 2019-03-25 Carl Kusche , Tom Reclik , Martina Freund , Talal Al-Samman , Ulrich Kerzel , Sandra Korte-Kerzel

Identifying multiple novel classes in an image, known as open-vocabulary multi-label recognition, is a challenging task in computer vision. Recent studies explore the transfer of powerful vision-language models such as CLIP. However, these…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Hao Tan , Zichang Tan , Jun Li , Ajian Liu , Jun Wan , Zhen Lei

Dual energy computed tomography (DECT) imaging plays an important role in advanced imaging applications due to its material decomposition capability. Image-domain decomposition operates directly on CT images using linear matrix inversion,…

Image and Video Processing · Electrical Eng. & Systems 2019-08-20 Zhipeng Li , Saiprasad Ravishankar , Yong Long , Jeffrey A. Fessler