English
Related papers

Related papers: Textual Inversion and Self-supervised Refinement f…

200 papers

A comprehensive understanding of vision and language and their interrelation are crucial to realize the underlying similarities and differences between these modalities and to learn more generalized, meaningful representations. In recent…

Computer Vision and Pattern Recognition · Computer Science 2021-12-10 Anindya Sundar Das , Sriparna Saha

Multi-image super-resolution (MISR) is a critical technique for satellite remote sensing. In the perspective of information, twin-image super-resolution (TISR) is regarded as the most challenging MISR scenario, having crucial applications…

Image and Video Processing · Electrical Eng. & Systems 2026-02-26 Chia-Hsiang Lin , Wei-Chih Liu , Yu-En Chiu , Jhao-Ting Lin

The automatic generation of radiology reports has emerged as a promising solution to reduce a time-consuming task and accurately capture critical disease-relevant findings in X-ray images. Previous approaches for radiology report generation…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Sang-Jun Park , Keun-Soo Heo , Dong-Hee Shin , Young-Han Son , Ji-Hye Oh , Tae-Eui Kam

In this paper, we introduce a new perspective for improving image restoration by removing degradation in the textual representations of a given degraded image. Intuitively, restoration is much easier on text modality than image one. For…

Computer Vision and Pattern Recognition · Computer Science 2024-01-01 Jingbo Lin , Zhilu Zhang , Yuxiang Wei , Dongwei Ren , Dongsheng Jiang , Wangmeng Zuo

Image restoration aims to recover degraded images. However, existing diffusion-based restoration methods, despite great success in natural image restoration, often struggle to faithfully reconstruct textual regions in degraded images. Those…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Jaewon Min , Jin Hyeon Kim , Paul Hyunbin Cho , Jaeeun Lee , Jihye Park , Minkyu Park , Sangpil Kim , Hyunhee Park , Seungryong Kim

Pre-training lays the foundation for recent successes in radiograph analysis supported by deep learning. It learns transferable image representations by conducting large-scale fully-supervised or self-supervised learning on a source domain.…

Image and Video Processing · Electrical Eng. & Systems 2022-01-28 Hong-Yu Zhou , Xiaoyu Chen , Yinghao Zhang , Ruibang Luo , Liansheng Wang , Yizhou Yu

Despite tremendous progress in computer vision, there has not been an attempt for machine learning on very large-scale medical image databases. We present an interleaved text/image deep learning system to extract and mine the semantic…

Computer Vision and Pattern Recognition · Computer Science 2015-05-05 Hoo-Chang Shin , Le Lu , Lauren Kim , Ari Seff , Jianhua Yao , Ronald M. Summers

Recently, generative diffusion priors have made huge strides as inverse problem solvers, including the ability to be adapted for inference on out-of-distribution data. Concurrently, implicit neural representations (INRs) have emerged as…

Image and Video Processing · Electrical Eng. & Systems 2026-03-12 Maliha Hossain , Haley Duba-Sullivan , Amirkoushyar Ziabari

The goal of scene text image super-resolution is to reconstruct high-resolution text-line images from unrecognizable low-resolution inputs. The existing methods relying on the optimization of pixel-level loss tend to yield text edges that…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Baolin Liu , Zongyuan Yang , Pengfei Wang , Junjie Zhou , Ziqi Liu , Ziyi Song , Yan Liu , Yongping Xiong

This paper addresses the task of interactive, conversational text-to-image retrieval. Our DIR-TIR framework progressively refines the target image search through two specialized modules: the Dialog Refiner Module and the Image Refiner…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Zongwei Zhen , Biqing Zeng

Multimodal Large Language Models (MLLMs) have shown strong potential for radiology report generation, yet their clinical translation is hindered by architectural heterogeneity and the prevalence of factual hallucinations. Standard…

Machine Learning · Computer Science 2026-01-13 Kun Zhao , Siyuan Dai , Pan Wang , Jifeng Song , Hui Ji , Chenghua Lin , Liang Zhan , Haoteng Tang

Radiology report generation aims at generating descriptive text from radiology images automatically, which may present an opportunity to improve radiology reporting and interpretation. A typical setting consists of training encoder-decoder…

Computation and Language · Computer Science 2021-09-28 An Yan , Zexue He , Xing Lu , Jiang Du , Eric Chang , Amilcare Gentili , Julian McAuley , Chun-Nan Hsu

Recent advances in deep learning have enabled researchers to explore tasks at the intersection of computer vision and natural language processing, such as image captioning, visual question answering, visual dialogue, and visual language…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Sonit Singh

Existing diffusion-based super-resolution approaches often exhibit semantic ambiguities due to inaccuracies and incompleteness in their text conditioning, coupled with the inherent tendency for cross-attention to divert towards irrelevant…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Chen Chen , Majid Abdolshah , Violetta Shevchenko , Hongdong Li , Chang Xu , Pulak Purkait

Automated recognition of texts in scenes has been a research challenge for years, largely due to the arbitrary variation of text appearances in perspective distortion, text line curvature, text styles and different types of imaging…

Computer Vision and Pattern Recognition · Computer Science 2019-04-03 Fangneng Zhan , Shijian Lu

Scene Text Image Super-Resolution (STISR) aims to enhance the resolution and legibility of text within low-resolution (LR) images, consequently elevating recognition accuracy in Scene Text Recognition (STR). Previous methods predominantly…

Computer Vision and Pattern Recognition · Computer Science 2023-11-23 Yuxuan Zhou , Liangcai Gao , Zhi Tang , Baole Wei

Textual Inversion (TI) is an efficient approach to text-to-image personalization but often fails on complex prompts. We trace these failures to embedding norm inflation: learned tokens drift to out-of-distribution magnitudes, degrading…

Machine Learning · Computer Science 2026-03-11 Kunhee Kim , NaHyeon Park , Kibeom Hong , Hyunjung Shim

Radiology reports are a rich resource for advancing deep learning applications in medicine by leveraging the large volume of data continuously being updated, integrated, and shared. However, there are significant challenges as well, largely…

Information Retrieval · Computer Science 2017-11-21 Imon Banerjee , Sriraman Madhavan , Roger Eric Goldman , Daniel L. Rubin

Recently, text-to-image (T2I) editing has been greatly pushed forward by applying diffusion models. Despite the visual promise of the generated images, inconsistencies with the expected textual prompt remain prevalent. This paper aims to…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Aoxue Li , Mingyang Yi , Zhenguo Li

Implicit neural representations (INRs) have emerged as a powerful paradigm for medical imaging via physics-informed unsupervised learning. Classical INRs optimize an entire network from scratch for each subject, leading to inefficient…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Qing Wu , Xuanyu Tian , Chenhe Du , Haonan Zhang , Xiao Wang , Le Lu , Yuyao Zhang