中文
相关论文

相关论文: CLNet: Cross-View Correspondence Makes a Stronger …

200 篇论文

Cross-modal alignment plays a crucial role in vision-language pre-training (VLP) models, enabling them to capture meaningful associations across different modalities. For this purpose, numerous masked modeling tasks have been proposed for…

计算机视觉与模式识别 · 计算机科学 2023-12-07 Rong-Cheng Tu , Yatai Ji , Jie Jiang , Weijie Kong , Chengfei Cai , Wenzhe Zhao , Hongfa Wang , Yujiu Yang , Wei Liu

Contrastive learning (CL) aims to preserve relational structure between samples by learning representations that reflect a similarity graph. Yet, the geometry of the resulting embeddings remains poorly understood. Here we show that weighted…

机器学习 · 计算机科学 2026-05-15 Raphael Vock , Edouard Duchesnay , Benoit Dufumier

Deep neural networks are playing an important role in state-of-the-art visual recognition. To represent high-level visual concepts, modern networks are equipped with large convolutional layers, which use a large number of filters and…

计算机视觉与模式识别 · 计算机科学 2017-03-06 Yan Wang , Lingxi Xie , Ya Zhang , Wenjun Zhang , Alan Yuille

Automatic and precise medical image segmentation (MIS) is of vital importance for clinical diagnosis and analysis. Current MIS methods mainly rely on the convolutional neural network (CNN) or self-attention mechanism (Transformer) for…

计算机视觉与模式识别 · 计算机科学 2024-08-28 Lanhu Wu , Miao Zhang , Yongri Piao , Zhenyan Yao , Weibing Sun , Feng Tian , Huchuan Lu

Modern learning-based visual feature extraction networks perform well in intra-domain localization, however, their performance significantly declines when image pairs are captured across long-term visual domain variations, such as different…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Zador Pataki , Mohammad Altillawi , Menelaos Kanakis , Rémi Pautrat , Fengyi Shen , Ziyuan Liu , Luc Van Gool , Marc Pollefeys

Unsupervised Multi-View Stereo (MVS) methods have achieved promising progress recently. However, previous methods primarily depend on the photometric consistency assumption, which may suffer from two limitations: indistinguishable regions…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Kaiqiang Xiong , Rui Peng , Zhe Zhang , Tianxing Feng , Jianbo Jiao , Feng Gao , Ronggang Wang

Few-shot segmentation aims to learn a segmentation model that can be generalized to novel classes with only a few training images. In this paper, we propose a Cross-Reference and Local-Global Conditional Networks (CRCNet) for few-shot…

计算机视觉与模式识别 · 计算机科学 2022-08-24 Weide Liu , Chi Zhang , Guosheng Lin , Fayao Liu

Aerial-ground localization is difficult due to large viewpoint and modality gaps between ground-level LiDAR and overhead imagery. We propose TransLocNet, a cross-modal attention framework that fuses LiDAR geometry with aerial semantic…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Phu Pham , Damon Conover , Aniket Bera

The existing human pose estimation methods are confronted with inaccurate long-distance regression or high computational cost due to the complex learning objectives. This work proposes a novel deep learning framework for human pose…

计算机视觉与模式识别 · 计算机科学 2021-05-18 ZiFan Chen , Xin Qin , Chao Yang , Li Zhang

Cross-view correspondence is a fundamental capability for spatial understanding and embodied AI. However, it is still far from being realized in Vision-Language Models (VLMs), especially in achieving precise point-level correspondence,…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Yipu Wang , Yuheng Ji , Yuyang Liu , Enshen Zhou , Ziqiang Yang , Yuxuan Tian , Ziheng Qin , Yue Liu , Huajie Tan , Cheng Chi , Zhiyuan Ma , Daniel Dajun Zeng , Xiaolong Zheng

ImageNet Large Scale Visual Recognition Challenge (ILSVRC) is one of the most authoritative academic competitions in the field of Computer Vision (CV) in recent years. But applying ILSVRC's annual champion directly to fine-grained visual…

计算机视觉与模式识别 · 计算机科学 2020-07-22 Fan Zhang , Meng Li , Guisheng Zhai , Yizhao Liu

Recovering the spatial layout of the cameras and the geometry of the scene from extreme-view images is a longstanding challenge in computer vision. Prevailing 3D reconstruction algorithms often adopt the image matching paradigm and presume…

计算机视觉与模式识别 · 计算机科学 2022-06-17 Wei-Chiu Ma , Anqi Joyce Yang , Shenlong Wang , Raquel Urtasun , Antonio Torralba

Multi-view clustering can explore common semantics from multiple views and has attracted increasing attention. However, existing works punish multiple objectives in the same feature space, where they ignore the conflict between learning…

机器学习 · 计算机科学 2022-03-28 Jie Xu , Huayi Tang , Yazhou Ren , Liang Peng , Xiaofeng Zhu , Lifang He

Convolutional neural networks (CNNs) are one of the most successful computer vision systems to solve object recognition. Furthermore, CNNs have major applications in understanding the nature of visual representations in the human brain. Yet…

计算机视觉与模式识别 · 计算机科学 2022-12-13 Amr Farahat , Felix Effenberger , Martin Vinck

Image transformation, a class of vision and graphics problems whose goal is to learn the mapping between an input image and an output image, develops rapidly in the context of deep neural networks. In Computer Vision (CV), many problems can…

计算机视觉与模式识别 · 计算机科学 2022-06-22 Yuanjie Yan , Suorong Yang , Yan Wang , Jian Zhao , Furao Shen

Although large-scale labeled data are essential for deep convolutional neural networks (ConvNets) to learn high-level semantic visual representations, it is time-consuming and impractical to collect and annotate large-scale datasets. A…

计算机视觉与模式识别 · 计算机科学 2023-10-06 Huili Huang , M. Mahdi Roozbahani

Cross-view image matching for geo-localisation is a challenging problem due to the significant visual difference between aerial and ground-level viewpoints. The method provides localisation capabilities from geo-referenced images,…

计算机视觉与模式识别 · 计算机科学 2024-09-25 Tavis Shore , Simon Hadfield , Oscar Mendez

Building extraction from remote sensing images is a challenging task due to the complex structure variations of the buildings. Existing methods employ convolutional or self-attention blocks to capture the multi-scale features in the…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Siyuan Yao , Dongxiu Liu , Taotao Li , Shengjie Li , Wenqi Ren , Xiaochun Cao

We present a simple, flexible, and general framework titled Partial Registration Network (PRNet), for partial-to-partial point cloud registration. Inspired by recently-proposed learning-based methods for registration, we use deep networks…

机器学习 · 计算机科学 2019-10-30 Yue Wang , Justin M. Solomon

Visual place recognition (VPR) is a highly challenging task that has a wide range of applications, including robot navigation and self-driving vehicles. VPR is particularly difficult due to the presence of duplicate regions and the lack of…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Yifan Xu , Pourya Shamsolmoali , Jie Yang