English

VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition

Computer Vision and Pattern Recognition 2025-12-30 v1 Artificial Intelligence

Abstract

Pedestrian Attribute Recognition (PAR) involves predicting fine-grained attributes such as clothing color, gender, and accessories from pedestrian imagery, yet is hindered by severe class imbalance, intricate attribute co-dependencies, and domain shifts. We introduce VLM-PAR, a modular vision-language framework built on frozen SigLIP 2 multilingual encoders. By first aligning image and prompt embeddings via refining visual features through a compact cross-attention fusion, VLM-PAR achieves significant accuracy improvement on the highly imbalanced PA100K benchmark, setting a new state-of-the-art performance, while also delivering significant gains in mean accuracy across PETA and Market-1501 benchmarks. These results underscore the efficacy of integrating large-scale vision-language pretraining with targeted cross-modal refinement to overcome imbalance and generalization challenges in PAR.

Keywords

Cite

@article{arxiv.2512.22217,
  title  = {VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition},
  author = {Abdellah Zakaria Sellam and Salah Eddine Bekhouche and Fadi Dornaika and Cosimo Distante and Abdenour Hadid},
  journal= {arXiv preprint arXiv:2512.22217},
  year   = {2025}
}
R2 v1 2026-07-01T08:41:55.542Z