English
Related papers

Related papers: FITA: Fine-grained Image-Text Aligner for Radiolog…

200 papers

Ultra-Wide-Field (UWF) retinal imaging has revolutionized retinal diagnostics by providing a comprehensive view of the retina. However, it often suffers from quality-degrading factors such as blurring and uneven illumination, which obscure…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Weicheng Liao , Zan Chen , Jianyang Xie , Yalin Zheng , Yuhui Ma , Yitian Zhao

Vision-language models (VLMs), such as CLIP and ALIGN, are generally trained on datasets consisting of image-caption pairs obtained from the web. However, real-world multimodal datasets, such as healthcare data, are significantly more…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Maya Varma , Jean-Benoit Delbrouck , Sarah Hooper , Akshay Chaudhari , Curtis Langlotz

E-commerce image search often takes a cropped image as the query, while each candidate is represented by full item images and structured text. This image-to-multimodal retrieval setting presents two asymmetries: a modality disparity -- a…

Information Retrieval · Computer Science 2026-05-19 Xinyu Sun , Huangyu Dai , Lingtao Mao , Zexin Zheng , Zihan Liang , Ben Chen , Chenyi Lei , Wenwu Ou

Computed Tomography (CT) based precise prostate segmentation for treatment planning is challenging due to (1) the unclear boundary of the prostate derived from CT's poor soft tissue contrast and (2) the limitation of convolutional neural…

Image and Video Processing · Electrical Eng. & Systems 2023-07-20 Chengyin Li , Yao Qiang , Rafi Ibn Sultan , Hassan Bagher-Ebadian , Prashant Khanduri , Indrin J. Chetty , Dongxiao Zhu

The field of text-conditioned image generation has made unparalleled progress with the recent advent of latent diffusion models. While remarkable, as the complexity of given text input increases, the state-of-the-art diffusion models may…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Jaskirat Singh , Liang Zheng

Fine-grained image recognition is very challenging due to the difficulty of capturing both semantic global features and discriminative local features. Meanwhile, these two features are not easy to be integrated, which are even conflicting…

Computer Vision and Pattern Recognition · Computer Science 2021-02-22 Shaokang Yang , Shuai Liu , Cheng Yang , Changhu Wang

Developing advanced medical imaging retrieval systems is challenging due to the varying definitions of `similar images' across different medical contexts. This challenge is compounded by the lack of large-scale, high-quality medical imaging…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Tengfei Zhang , Ziheng Zhao , Chaoyi Wu , Xiao Zhou , Ya Zhang , Yanfeng Wang , Weidi Xie

Integrating multimodal knowledge for abstractive summarization task is a work-in-progress research area, with present techniques inheriting fusion-then-generation paradigm. Due to semantic gaps between computer vision and natural language…

Artificial Intelligence · Computer Science 2022-08-09 Zijian Zhang , Chang Shu , Youxin Chen , Jing Xiao , Qian Zhang , Lu Zheng

Remote sensing (RS) image-text retrieval plays a critical role in understanding massive RS imagery. However, the dense multi-object distribution and complex backgrounds in RS imagery make it difficult to simultaneously achieve fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Xi Chen , Xu Chen , Xiangyang Jia , Xu Zhang , Shuquan Wei , Wei Wang

TIReID aims to retrieve the image corresponding to the given text query from a pool of candidate images. Existing methods employ prior knowledge from single-modality pre-training to facilitate learning, but lack multi-modal correspondences.…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Shuanglin Yan , Neng Dong , Liyan Zhang , Jinhui Tang

Recent advances in text-to-video diffusion models have enabled high-quality video synthesis, but controllable generation remains challenging, particularly under limited data and compute. Existing fine-tuning methods for conditional…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Kinam Kim , Junha Hyung , Jaegul Choo

Despite recent advances in deep generative modeling, skin lesion classification systems remain constrained by the limited availability of large, diverse, and well-annotated clinical datasets, resulting in class imbalance between benign and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Stathis Galanakis , Alexandros Koliousis , Stefanos Zafeiriou

Text-to-video generation has shown promising results. However, by taking only natural languages as input, users often face difficulties in providing detailed information to precisely control the model's output. In this work, we propose…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Hsin-Ping Huang , Yu-Chuan Su , Deqing Sun , Lu Jiang , Xuhui Jia , Yukun Zhu , Ming-Hsuan Yang

With the advent of large pre-trained transformer models, fine-tuning these models for various downstream tasks is a critical problem. Paucity of training data, the existence of data silos, and stringent privacy constraints exacerbate this…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Naif Alkhunaizi , Faris Almalik , Rouqaiah Al-Refai , Muzammal Naseer , Karthik Nandakumar

Current full-reference image quality assessment (FR-IQA) methods often fuse features from reference and distorted images, overlooking that color and luminance distortions occur mainly at low frequencies, whereas edge and texture distortions…

Image and Video Processing · Electrical Eng. & Systems 2024-12-23 Xuekai Wei , Junyu Zhang , Qinlin Hu , Mingliang Zhou\\Yong Feng , Weizhi Xian , Huayan Pu , Sam Kwong

Nowadays, a huge number of images are available. However, retrieving a required image for an ordinary user is a challenging task in computer vision systems. During the past two decades, many types of research have been introduced to improve…

Multimedia · Computer Science 2020-01-30 Amir Vatani , Milad Taleby Ahvanooey , Mostafa Rahimi

Data scarcity and privacy concerns limit the availability of high-quality medical images for public use, which can be mitigated through medical image synthesis. However, current medical image synthesis methods often struggle to accurately…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Wenting Chen , Pengyu Wang , Hui Ren , Lichao Sun , Quanzheng Li , Yixuan Yuan , Xiang Li

Remote Sensing Image-Text Retrieval (RSITR) is pivotal for knowledge services and data mining in the remote sensing (RS) domain. Considering the multi-scale representations in image content and text vocabulary can enable the models to learn…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Rui Yang , Shuang Wang , Yingping Han , Yuanheng Li , Dong Zhao , Dou Quan , Yanhe Guo , Licheng Jiao

We propose a novel implicit feature refinement module for high-quality instance segmentation. Existing image/video instance segmentation methods rely on explicitly stacked convolutions to refine instance features before the final…

Computer Vision and Pattern Recognition · Computer Science 2021-12-10 Lufan Ma , Tiancai Wang , Bin Dong , Jiangpeng Yan , Xiu Li , Xiangyu Zhang

We present Autoregressive Representation Alignment (ARRA), a new training framework that unlocks global-coherent text-to-image generation in autoregressive LLMs without architectural modifications. Different from prior works that require…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Xing Xie , Jiawei Liu , Ziyue Lin , Huijie Fan , Zhi Han , Yandong Tang , Liangqiong Qu