English

TKwinFormer: Top k Window Attention in Vision Transformers for Feature Matching

Image and Video Processing 2025-04-01 v2

Abstract

Local feature matching remains a challenging task, primarily due to difficulties in matching sparse keypoints and low-texture regions. The key to solving this problem lies in effectively and accurately integrating global and local information. To achieve this goal, we introduce an innovative local feature matching method called TKwinFormer. Our approach employs a multi-stage matching strategy to optimize the efficiency of information interaction. Furthermore, we propose a novel attention mechanism called Top K Window Attention, which facilitates global information interaction through window tokens prior to patch-level matching, resulting in improved matching accuracy. Additionally, we design an attention block to enhance attention between channels. Experimental results demonstrate that TKwinFormer outperforms state-of-the-art methods on various benchmarks. Code is available at: https://github.com/LiaoYun0x0/TKwinFormer.

Keywords

Cite

@article{arxiv.2308.15144,
  title  = {TKwinFormer: Top k Window Attention in Vision Transformers for Feature Matching},
  author = {Yun Liao and Yide Di and Hao Zhou and Kaijun Zhu and Mingyu Lu and Yijia Zhang and Qing Duan and Junhui Liu},
  journal= {arXiv preprint arXiv:2308.15144},
  year   = {2025}
}

Comments

After careful reconsideration, we have decided to withdraw the manuscript due to data inconsistencies and issues with methodology. Given these concerns, we believe it would be inappropriate to proceed with the revised version, and we have therefore decided to retract our submission

R2 v1 2026-06-28T12:07:06.817Z