English

Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis

Computer Vision and Pattern Recognition 2025-10-14 v1

Abstract

Accurate traffic congestion classification is essential for intelligent transportation systems and real-time urban traffic management. This paper presents a multimodal framework combining open-vocabulary visual-language reasoning (CLIP), object detection (YOLO-World), and motion analysis via MOG2-based background subtraction. The system predicts congestion levels on an ordinal scale from 1 (free flow) to 5 (severe congestion), enabling semantically aligned and temporally consistent classification. To enhance interpretability, we incorporate motion-based confidence weighting and generate annotated visual outputs. Experimental results show the model achieves 76.7 percent accuracy, an F1 score of 0.752, and a Quadratic Weighted Kappa (QWK) of 0.684, significantly outperforming unimodal baselines. These results demonstrate the framework's effectiveness in preserving ordinal structure and leveraging visual-language and motion modalities. Future enhancements include incorporating vehicle sizing and refined density metrics.

Keywords

Cite

@article{arxiv.2510.10342,
  title  = {Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis},
  author = {Yu-Hsuan Lin},
  journal= {arXiv preprint arXiv:2510.10342},
  year   = {2025}
}

Comments

7 pages, 4 figures. Preprint submitted to arXiv in October 2025

R2 v1 2026-07-01T06:31:44.268Z