English

Walking Further: Semantic-aware Multimodal Gait Recognition Under Long-Range Conditions

Computer Vision and Pattern Recognition 2026-03-17 v1 Artificial Intelligence

Abstract

Gait recognition is an emerging biometric technology that enables non-intrusive and hard-to-spoof human identification. However, most existing methods are confined to short-range, unimodal settings and fail to generalize to long-range and cross-distance scenarios under real-world conditions. To address this gap, we present \textbf{LRGait}, the first LiDAR-Camera multimodal benchmark designed for robust long-range gait recognition across diverse outdoor distances and environments. We further propose \textbf{EMGaitNet}, an end-to-end framework tailored for long-range multimodal gait recognition. To bridge the modality gap between RGB images and point clouds, we introduce a semantic-guided fusion pipeline. A CLIP-based Semantic Mining (SeMi) module first extracts human body-part-aware semantic cues, which are then employed to align 2D and 3D features via a Semantic-Guided Alignment (SGA) module within a unified embedding space. A Symmetric Cross-Attention Fusion (SCAF) module hierarchically integrates visual contours and 3D geometric features, and a Spatio-Temporal (ST) module captures global gait dynamics. Extensive experiments on various gait datasets validate the effectiveness of our method.

Keywords

Cite

@article{arxiv.2603.14189,
  title  = {Walking Further: Semantic-aware Multimodal Gait Recognition Under Long-Range Conditions},
  author = {Zhiyang Lu and Wen Jiang and Tianren Wu and Zhichao Wang and Changwang Zhang and Siqi Shen and Ming Cheng},
  journal= {arXiv preprint arXiv:2603.14189},
  year   = {2026}
}

Comments

Accepted by AAAI 2026

R2 v1 2026-07-01T11:20:27.603Z