中文
相关论文

相关论文: LoFi: Location-Aware Fine-Grained Representation L…

200 篇论文

Recent work in interpretability shows that large language models (LLMs) can be adapted for new tasks in a learning-free way: it is possible to intervene on LLM representations to elicit desired behaviors for alignment. For instance, adding…

计算与语言 · 计算机科学 2024-11-01 Fangcong Yin , Xi Ye , Greg Durrett

The impressive performance of Large Language Model (LLM) has prompted researchers to develop Multi-modal LLM (MLLM), which has shown great potential for various multi-modal tasks. However, current MLLM often struggles to effectively address…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Yeyuan Wang , Dehong Gao , Bin Li , Rujiao Long , Lei Yi , Xiaoyan Cai , Libin Yang , Jinxia Zhang , Shanqing Yu , Qi Xuan

Differential medical VQA models compare multiple images to identify clinically meaningful changes and rely on vision encoders to capture fine-grained visual differences that reflect radiologists' comparative diagnostic workflows. However,…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Denis Musinguzi , Caren Han , Prasenjit Mitra

The lack of fine-grained annotations hinders the deployment of automated diagnosis systems, which require human-interpretable justification for their decision process. In this paper, we address the problem of weakly supervised…

计算机视觉与模式识别 · 计算机科学 2022-10-10 Constantin Seibold , Jens Kleesiek , Heinz-Peter Schlemmer , Rainer Stiefelhagen

Neural fields or implicit neural representations (INRs) have attracted significant attention in computer vision and imaging due to their efficient coordinate-based representation of images and 3D volumes. In this work, we introduce a…

计算机视觉与模式识别 · 计算机科学 2024-12-24 AmirEhsan Khorashadizadeh , Tobías I. Liaudat , Tianlin Liu , Jason D. McEwen , Ivan Dokmanić

The scarcity of richly annotated medical images is limiting supervised deep learning based solutions to medical image analysis tasks, such as localizing discriminatory radiomic disease signatures. Therefore, it is desirable to leverage…

计算机视觉与模式识别 · 计算机科学 2019-06-10 Saeid Asgari Taghanaki , Mohammad Havaei , Tess Berthier , Francis Dutil , Lisa Di Jorio , Ghassan Hamarneh , Yoshua Bengio

The self-supervised contrastive learning strategy has attracted considerable attention due to its exceptional ability in representation learning. However, current contrastive learning tends to learn global coarse-grained representations of…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Jialu Shi , Zhiqiang Wei , Jie Nie , Lei Huang

Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Jiale Liu , Haoming Zhou , Yishu Liu , Bingzhi Chen , Yuncheng Jiang

Fine-tuning adapts pretrained models for specific tasks but poses the risk of catastrophic forgetting (CF), where critical knowledge from pretraining is overwritten. To address the issue of CF in a general-purpose framework, we propose…

计算与语言 · 计算机科学 2025-11-25 Runyu Wang , Peng Ping , Zhengyu Guo , Xiaoye Zhang , Quan Shi , Liting Zhou , Tianbo Ji

This paper tackles the problem of learning a finer representation than the one provided by training labels. This enables fine-grained category retrieval of images in a collection annotated with coarse labels only. Our network is learned…

计算机视觉与模式识别 · 计算机科学 2020-11-30 Hugo Touvron , Alexandre Sablayrolles , Matthijs Douze , Matthieu Cord , Hervé Jégou

Understanding how deep neural networks learn useful internal representations from data remains a central open problem in the theory of deep learning. We introduce Neural Low-Degree Filtering (Neural LoFi), a stylized limit of gradient-based…

机器学习 · 计算机科学 2026-05-14 Yatin Dandi , Matteo Vilucchio , Luca Arnaboldi , Hugo Tabanelli , Florent Krzakala

Multimodal automatic speech recognition systems integrate information from images to improve speech recognition quality, by grounding the speech in the visual context. While visual signals have been shown to be useful for recovering…

计算与语言 · 计算机科学 2020-10-07 Tejas Srinivasan , Ramon Sanabria , Florian Metze , Desmond Elliott

Chest X-ray imaging is commonly used to diagnose pneumonia, but accurately localizing the pneumonia-affected regions typically requires detailed pixel-level annotations, which are costly and time consuming to obtain. To address this…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Kiran Shahi , Anup Bagale

Federated learning enables building a shared model from multicentre data while storing the training data locally for privacy. In this paper, we present an evaluation (called CXR-FL) of deep learning-based models for chest X-ray image…

图像与视频处理 · 电气工程与系统科学 2022-08-09 Filip Ślazyk , Przemysław Jabłecki , Aneta Lisowska , Maciej Malawski , Szymon Płotka

Chest X-rays have become the focus of vigorous deep learning research in recent years due to the availability of large labeled datasets. While classification of anomalous findings is now possible, ensuring that they are correctly localized…

图像与视频处理 · 电气工程与系统科学 2022-04-22 Neha Srivathsa , Razi Mahmood , Tanveer Syeda-Mahmood

Despite significant progress in talking head synthesis since the introduction of Neural Radiance Fields (NeRF), visual artifacts and high training costs persist as major obstacles to large-scale commercial adoption. We propose that…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Tianqi Li , Ruobing Zheng , Bonan Li , Zicheng Zhang , Meng Wang , Jingdong Chen , Ming Yang

Chest radiographs are the most common diagnostic exam in emergency rooms and intensive care units today. Recently, a number of researchers have begun working on large chest X-ray datasets to develop deep learning models for recognition of a…

计算机视觉与模式识别 · 计算机科学 2020-11-20 Tanveer Syeda-Mahmood , Ph. D , K. C. L Wong , Ph. D , Joy T. Wu , M. D. , M. P. H , Ashutosh Jadhav , Ph. D , Orest Boyko , M. D. Ph. D

Fine-grained image recognition is very challenging due to the difficulty of capturing both semantic global features and discriminative local features. Meanwhile, these two features are not easy to be integrated, which are even conflicting…

计算机视觉与模式识别 · 计算机科学 2021-02-22 Shaokang Yang , Shuai Liu , Cheng Yang , Changhu Wang

Given a ground-level query image and a geo-referenced aerial image that covers the query's local surroundings, fine-grained cross-view localization aims to estimate the location of the ground camera inside the aerial image. Recent works…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Zimin Xia , Yujiao Shi , Hongdong Li , Julian F. P. Kooij

Long-term visual localization is the problem of estimating the camera pose of a given query image in a scene whose appearance changes over time. It is an important problem in practice, for example, encountered in autonomous driving. In…

计算机视觉与模式识别 · 计算机科学 2019-08-20 Måns Larsson , Erik Stenborg , Carl Toft , Lars Hammarstrand , Torsten Sattler , Fredrik Kahl
‹ 上一页 1 2 3 10 下一页 ›