English
Related papers

Related papers: Prototype-Based Knowledge Guidance for Fine-Graine…

200 papers

Radiology report generation aims to produce computer-aided diagnoses to alleviate the workload of radiologists and has drawn increasing attention recently. However, previous deep learning methods tend to neglect the mutual influences…

Computation and Language · Computer Science 2022-01-12 Song Wang , Liyan Tang , Mingquan Lin , George Shih , Ying Ding , Yifan Peng

Structured reconstruction is a non-trivial dense prediction problem, which extracts structural information (\eg, building corners and edges) from a raster image, then reconstructs it to a 2D planar graph accordingly. Compared with common…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Hongbo Tian , Yulong Li , Linzhi Huang , Xu Ling , Yue Yang , Jiani Hu

Self-supervised learning (SSL) has emerged as a powerful technique for learning visual representations. While recent SSL approaches achieve strong results in global image understanding, they are limited in capturing the structured…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Oussama Hadjerci , Antoine Letienne , Mohamed Abbas Hedjazi , Adel Hafiane

The current gold standard for evaluating generated chest x-ray (CXR) reports is through radiologist annotations. However, this process can be extremely time-consuming and costly, especially when evaluating large numbers of reports. In this…

Computation and Language · Computer Science 2024-08-13 Alyssa Huang , Oishi Banerjee , Kay Wu , Eduardo Pontes Reis , Pranav Rajpurkar

Subject-driven text-to-image generation still struggles to preserve high-frequency identity details such as logos, patterns, and text. Existing methods typically operate directly in RGB space, which often leads to detail degradation under…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Hanzhong Guo , Yizhou Yu

Scanning transmission electron microscopy is a common tool used to study the atomic structure of materials. It is an inherently multimodal tool allowing for the simultaneous acquisition of multiple information channels. Despite its…

Large-scale contrastive pre-training produces powerful Vision-and-Language Models (VLMs) capable of generating representations (embeddings) effective for a wide variety of visual and multimodal tasks. However, these pretrained embeddings…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Nikolaos-Antonios Ypsilantis , Kaifeng Chen , André Araujo , Ondřej Chum

Retrieval-Augmented Generation (RAG) helps large language models (LLMs) answer knowledge-intensive and time-sensitive questions by conditioning generation on external evidence. However, most RAG systems still retrieve unstructured chunks…

Computation and Language · Computer Science 2026-03-11 Jiashuo Sun , Yixuan Xie , Jimeng Shi , Shaowen Wang , Jiawei Han

Vertebral fracture grading classifies the severity of vertebral fractures, which is a challenging task in medical imaging and has recently attracted Deep Learning (DL) models. Only a few works attempted to make such models…

Image registration is an ill-posed dense vision task, where multiple solutions achieve similar loss values, motivating probabilistic inference. Variational inference has previously been employed to capture these distributions, however…

Image and Video Processing · Electrical Eng. & Systems 2026-03-19 Ivor J. A. Simpson , Neill D. F. Campbell

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for integrating spatial cues, such as point clouds or depth, either…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Yang Liu , Ming Ma , Xiaomin Yu , Pengxiang Ding , Han Zhao , Mingyang Sun , Siteng Huang , Donglin Wang

Diffusion magnetic resonance imaging (dMRI) is a crucial non-invasive technique for exploring the microstructure of the living human brain. Traditional hand-crafted and model-based tissue microstructure reconstruction methods often require…

Image and Video Processing · Electrical Eng. & Systems 2025-02-26 Xinrui Ma , Jian Cheng , Wenxin Fan , Ruoyou Wu , Yongquan Ye , Shanshan Wang

High-resolution (HR) land-cover mapping is often constrained by the high cost of dense HR annotations. We revisit this problem from the perspective of map super-resolution, which enhances coarse low-resolution (LR) land-cover products into…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Ruiqi Wang , Qi Yu , Jie Ma , Hanlin Wu

Image Super-Resolution (SR) aims to reconstruct high-resolution images from degraded low-resolution inputs. While diffusion-based SR methods offer powerful generative capabilities, their performance heavily depends on how semantic priors…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Lei Jiang , Xin Liu , Xinze Tong , Zhiliang Li , Jie Liu , Jie Tang , Gangshan Wu

The approximation and convergence properties of implicit neural representations (INRs) are known to be highly sensitive to parameter initialization strategies. While several data-driven initialization methods demonstrate significant…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Kushal Vyas , Alper Kayabasi , Daniel Kim , Vishwanath Saragadam , Ashok Veeraraghavan , Guha Balakrishnan

We present Pre-trained Machine Reader (PMR), a novel method for retrofitting pre-trained masked language models (MLMs) to pre-trained machine reading comprehension (MRC) models without acquiring labeled data. PMR can resolve the discrepancy…

Computation and Language · Computer Science 2023-10-17 Weiwen Xu , Xin Li , Wenxuan Zhang , Meng Zhou , Wai Lam , Luo Si , Lidong Bing

Training large, general-purpose language models poses significant challenges. The growing availability of specialized expert models, fine-tuned from pretrained models for specific tasks or domains, offers a promising alternative. Leveraging…

Computation and Language · Computer Science 2025-08-19 William Fleshman , Benjamin Van Durme

Multi-modality medical imaging is crucial in clinical treatment as it can provide complementary information for medical image segmentation. However, collecting multi-modal data in clinical is difficult due to the limitation of the scan time…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Shuai Wang , Zipei Yan , Daoan Zhang , Haining Wei , Zhongsen Li , Rui Li

Scene text recognition has witnessed rapid development with the advance of convolutional neural networks. Nonetheless, most of the previous methods may not work well in recognizing text with low resolution which is often seen in natural…

Computer Vision and Pattern Recognition · Computer Science 2019-10-22 Wenjia Wang , Enze Xie , Peize Sun , Wenhai Wang , Lixun Tian , Chunhua Shen , Ping Luo

Large Language Models (LLMs) have demonstrated exceptional performance in natural language processing tasks, yet their massive size makes serving them inefficient and costly. Semi-structured pruning has emerged as an effective method for…

Machine Learning · Computer Science 2025-06-25 Hongyi Liu , Rajarshi Saha , Zhen Jia , Youngsuk Park , Jiaji Huang , Shoham Sabach , Yu-Xiang Wang , George Karypis