English
Related papers

Related papers: Textual Inversion and Self-supervised Refinement f…

200 papers

Automatically generating medical reports for retinal images is one of the promising ways to help ophthalmologists reduce their workload and improve work efficiency. In this work, we propose a new context-driven encoding network to…

Computer Vision and Pattern Recognition · Computer Science 2021-06-01 Jia-Hong Huang , Ting-Wei Wu , Chao-Han Huck Yang , Marcel Worring

Supervised Deep-Learning (DL)-based reconstruction algorithms have shown state-of-the-art results for highly-undersampled dynamic Magnetic Resonance Imaging (MRI) reconstruction. However, the requirement of excessive high-quality…

Image and Video Processing · Electrical Eng. & Systems 2025-11-11 Jie Feng , Ruimin Feng , Qing Wu , Zhiyong Zhang , Yuyao Zhang , Hongjiang Wei

In this research, we introduce RefineNet, a novel architecture designed to address resolution limitations in text-to-image conversion systems. We explore the challenges of generating high-resolution images from textual descriptions,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-01 Fan Shi

Multiview super-resolution image reconstruction (SRIR) is often cast as a resampling problem by merging non-redundant data from multiple low-resolution (LR) images on a finer high-resolution (HR) grid, while inverting the effect of the…

Computer Vision and Pattern Recognition · Computer Science 2017-05-04 Vildan Atalay Aydin , Hassan Foroosh

We present Transformation Invariance and Covariance Contrast (TiCo) for self-supervised visual representation learning. Similar to other recent self-supervised learning methods, our method is based on maximizing the agreement among…

Computer Vision and Pattern Recognition · Computer Science 2022-06-24 Jiachen Zhu , Rafael M. Moraes , Serkan Karakulak , Vlad Sobol , Alfredo Canziani , Yann LeCun

Diffusion models have demonstrated impressive performance in text-guided image generation. Current methods that leverage the knowledge of these models for image editing either fine-tune them using the input image (e.g., Imagic) or…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Zhongping Zhang , Jian Zheng , Jacob Zhiyuan Fang , Bryan A. Plummer

Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image quality and textual readability. To address this, we introduce TIGER (Text-Image Guided…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Minxing Luo , Linlong Fan , Wang Qiushi , Ge Wu , Yiyan Luo , Yuhang Yu , Jinwei Chen , Yaxing Wang , Qingnan Fan , Jian Yang

We study the interpolation capabilities of implicit neural representations (INRs) of images. In principle, INRs promise a number of advantages, such as continuous derivatives and arbitrary sampling, being freed from the restrictions of a…

Image and Video Processing · Electrical Eng. & Systems 2024-05-02 Lorenzo Luzi , Daniel LeJeune , Ali Siahkoohi , Sina Alemohammad , Vishwanath Saragadam , Hossein Babaei , Naiming Liu , Zichao Wang , Richard G. Baraniuk

We propose a cross-modal transformer-based neural correction models that refines the output of an automatic speech recognition (ASR) system so as to exclude ASR errors. Generally, neural correction models are composed of encoder-decoder…

Computation and Language · Computer Science 2021-07-06 Tomohiro Tanaka , Ryo Masumura , Mana Ihori , Akihiko Takashima , Takafumi Moriya , Takanori Ashihara , Shota Orihashi , Naoki Makishima

The goal of self-supervised learning from images is to construct image representations that are semantically meaningful via pretext tasks that do not require semantic annotations for a large training set of images. Many pretext tasks lead…

Computer Vision and Pattern Recognition · Computer Science 2019-12-05 Ishan Misra , Laurens van der Maaten

With the rise of large, publicly-available text-to-image diffusion models, text-guided real image editing has garnered much research attention recently. Existing methods tend to either rely on some form of per-instance or per-task…

Computer Vision and Pattern Recognition · Computer Science 2022-11-16 Adham Elarabawy , Harish Kamath , Samuel Denton

Single Image Super-Resolution (SISR) aims to generate a high-resolution (HR) image of a given low-resolution (LR) image. The most of existing convolutional neural network (CNN) based SISR methods usually take an assumption that a LR image…

Image and Video Processing · Electrical Eng. & Systems 2019-09-10 Rao Muhammad Umer , Gian Luca Foresti , Christian Micheloni

Writing reports by analyzing medical images is error-prone for inexperienced practitioners and time consuming for experienced ones. In this work, we present RepsNet that adapts pre-trained vision and language models to interpret medical…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Ajay Kumar Tanwani , Joelle Barral , Daniel Freedman

Language-supervised pre-training has proven to be a valuable method for extracting semantically meaningful features from images, serving as a foundational element in multimodal systems within the computer vision and medical imaging domains.…

Image Super-Resolution (SR) aims to reconstruct high-resolution images from degraded low-resolution inputs. While diffusion-based SR methods offer powerful generative capabilities, their performance heavily depends on how semantic priors…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Lei Jiang , Xin Liu , Xinze Tong , Zhiliang Li , Jie Liu , Jie Tang , Gangshan Wu

Text-to-image generation models~(e.g., Stable Diffusion) have achieved significant advancements, enabling the creation of high-quality and realistic images based on textual descriptions. Prompt inversion, the task of identifying the textual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Mingzhe Li , Kejing Xia , Gehao Zhang , Zhenting Wang , Guanhong Tao , Siqi Pan , Juan Zhai , Shiqing Ma

Existing Referring Image Segmentation (RIS) methods typically require expensive pixel-level or box-level annotations for supervision. In this paper, we observe that the referring texts used in RIS already provide sufficient information to…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Fang Liu , Yuhao Liu , Yuqiu Kong , Ke Xu , Lihe Zhang , Baocai Yin , Gerhard Hancke , Rynson Lau

State of the art magnetic resonance (MR) image super-resolution methods (ISR) using convolutional neural networks (CNNs) leverage limited contextual information due to the limited spatial coverage of CNNs. Vision transformers (ViT) learn…

Image and Video Processing · Electrical Eng. & Systems 2022-07-26 Dwarikanath Mahapatra

Radiology report generation (RRG) aims to automatically produce clinically accurate textual reports from medical images. Existing methods predominantly rely on autoregressive (AR) language models, whose causal dependency structure restricts…

Artificial Intelligence · Computer Science 2026-05-19 Shiying Yu , Jielei Wang , Guoming Lu

Composed Image Retrieval (CIR) aims to retrieve a target image based on a reference image and conditioning text, enabling controllable image searches. The mainstream Zero-Shot (ZS) CIR methods bypass the need for expensive training CIR…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Jaeseok Byun , Seokhyeon Jeong , Wonjae Kim , Sanghyuk Chun , Taesup Moon
‹ Prev 1 3 4 5 6 7 10 Next ›