English
Related papers

Related papers: LoFi: Location-Aware Fine-Grained Representation L…

200 papers

Fine-grained recognition is challenging due to its subtle local inter-class differences versus large intra-class variations such as poses. A key to address this problem is to localize discriminative parts to extract pose-invariant features.…

Computer Vision and Pattern Recognition · Computer Science 2017-03-22 Xiao Liu , Tian Xia , Jiang Wang , Yi Yang , Feng Zhou , Yuanqing Lin

Despite tremendous efforts, it is very challenging to generate a robust model to assist in the accurate quantification assessment of COVID-19 on chest CT images. Due to the nature of blurred boundaries, the supervised segmentation methods…

Image and Video Processing · Electrical Eng. & Systems 2021-03-02 Yang Yang , Jiancong Chen , Ruixuan Wang , Ting Ma , Lingwei Wang , Jie Chen , Wei-Shi Zheng , Tong Zhang

Low-rank adaptation (LoRA) has become a prevalent method for adapting pre-trained large language models to downstream tasks. However, the simple low-rank decomposition form may constrain the hypothesis space. To address this limitation, we…

Machine Learning · Computer Science 2025-04-30 Zhekai Du , Yinjie Min , Jingjing Li , Ke Lu , Changliang Zou , Liuhua Peng , Tingjin Chu , Mingming Gong

Pre-trained large-scale language models (LLMs) excel at producing coherent articles, yet their outputs may be untruthful, toxic, or fail to align with user expectations. Current approaches focus on using reinforcement learning with human…

Computation and Language · Computer Science 2024-06-06 Dehong Xu , Liang Qiu , Minseok Kim , Faisal Ladhak , Jaeyoung Do

Protein function is inherently linked to its localization within the cell, and fluorescent microscopy data is an indispensable resource for learning representations of proteins. Despite major developments in molecular representation…

Quantitative Methods · Quantitative Biology 2022-05-25 Anastasia Razdaibiedina , Alexander Brechalov

Despite significant advancements in License Plate Recognition (LPR) through deep learning, most improvements rely on high-resolution images with clear characters. This scenario does not reflect real-world conditions where traffic…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Valfride Nascimento , Rayson Laroca , Rafael O. Ribeiro , William Robson Schwartz , David Menotti

Recent advancements in multimodal Large Language Models (LLMs) have significantly enhanced the automation of medical image analysis, particularly in generating radiology reports from chest X-rays (CXR). However, these models still suffer…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Yunsoo Kim , Jinge Wu , Su-Hwan Kim , Pardeep Vasudev , Jiashu Shen , Honghan Wu

Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease. While such capability is largely attributed to the rich world…

Computer Vision and Pattern Recognition · Computer Science 2025-01-08 Jeonghwan Kim , Heng Ji

Efficient vision-language understanding of large Remote Sensing Images (RSIs) is meaningful but challenging. Current Large Vision-Language Models (LVLMs) typically employ limited pre-defined grids to process images, leading to information…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Junwei Luo , Yingying Zhang , Xue Yang , Kang Wu , Qi Zhu , Lei Liang , Jingdong Chen , Yansheng Li

Place recognition is a challenging but crucial task in robotics. Current description-based methods may be limited by representation capabilities, while pairwise similarity-based methods require exhaustive searches, which is time-consuming.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Chencan Fu , Lin Li , Jianbiao Mei , Yukai Ma , Linpeng Peng , Xiangrui Zhao , Yong Liu

Report generation models offer fine-grained textual interpretations of medical images like chest X-rays, yet they often lack interactivity (i.e. the ability to steer the generation process through user queries) and localized…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Philip Müller , Georgios Kaissis , Daniel Rueckert

Precise modeling of lane topology is essential for autonomous driving, as it directly impacts navigation and control decisions. Existing methods typically represent each lane with a single query and infer topological connectivity based on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Guoqing Xu , Yiheng Li , Yang Yang

Large Language Models (LLMs) have shown strong performance in text-based healthcare tasks. However, their utility in image-based applications remains unexplored. We investigate the effectiveness of LLMs for medical imaging tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Felicia Liu , Jay J. Yoo , Farzad Khalvati

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains, particularly in task generalization for both text and vision data. While fine-tuning these models can significantly enhance their performance on…

Machine Learning · Computer Science 2025-01-15 Navyansh Mahla , Kshitij Sharad Jadhav , Ganesh Ramakrishnan

Fine-grained recognition involves the classification of images from subordinate macro-categories, and it is challenging due to small inter-class differences. To overcome this, most methods perform discriminative feature selection enabled by…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Edwin Arkel Rios , Min-Chun Hu , Bo-Cheng Lai

Lip-reading aims to recognize speech content from videos via visual analysis of speakers' lip movements. This is a challenging task due to the existence of homophemes-words which involve identical or highly similar lip movements, as well as…

Computer Vision and Pattern Recognition · Computer Science 2019-09-04 Chenhao Wang

The automatic detection of critical findings in chest X-rays (CXR), such as pneumothorax, is important for assisting radiologists in their clinical workflow like triaging time-sensitive cases and screening for incidental findings. While…

Machine Learning · Computer Science 2020-01-27 Evan Schwab , André Gooßen , Hrishikesh Deshpande , Axel Saalbach

Fine-Grained Visual Recognition (FGVR) tackles the problem of distinguishing highly similar categories. One of the main approaches to FGVR, namely subset learning, tries to leverage information from existing class taxonomies to improve the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Pablo Villacorta , Jesús M. Rodríguez-de-Vera , Marc Bolaños , Ignacio Sarasúa , Bhalaji Nagarajan , Petia Radeva

In this work we focus on learning facial representations that can be adapted to train effective face recognition models, particularly in the absence of labels. Firstly, compared with existing labelled face datasets, a vastly larger…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Zhonglin Sun , Chen Feng , Ioannis Patras , Georgios Tzimiropoulos

Recently, both Contrastive Learning (CL) and Mask Image Modeling (MIM) demonstrate that self-supervision is powerful to learn good representations. However, naively combining them is far from success. In this paper, we start by making the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-01 Ziyu Jiang , Yinpeng Chen , Mengchen Liu , Dongdong Chen , Xiyang Dai , Lu Yuan , Zicheng Liu , Zhangyang Wang