English

Benchmarking Video Foundation Models for Remote Parkinson's Disease Screening

Computer Vision and Pattern Recognition 2026-02-27 v2

Abstract

Video-based assessments offer a scalable pathway for remote Parkinson's disease (PD) screening. While traditional approaches rely on handcrafted features mimicking clinical scales, recent advances in video foundation models (VFMs) enable representation learning without task-specific customization. However, the comparative effectiveness of different VFM architectures across diverse clinical tasks remains poorly understood. We present a large-scale systematic study using a novel video dataset from 1,888 participants (727 with PD), comprising 32,847 videos across 16 standardized clinical tasks. We evaluate seven state-of-the-art VFMs -- including VideoPrism, V-JEPA, ViViT, and VideoMAE -- to determine their robustness in clinical screening. By evaluating frozen embeddings with a linear classification head, we demonstrate that task saliency is highly model-dependent: VideoPrism excels in capturing visual speech kinematics (no audio) and facial expressivity, while V-JEPA proves superior for upper-limb motor tasks. Notably, TimeSformer remains highly competitive for rhythmic tasks like finger tapping. Our experiments yield AUCs of 76.4 - 85.3% and accuracies of 71.5 - 80.6%. While high specificity (up to 90.3%) suggests strong potential for ruling out healthy individuals, the lower sensitivity (43.2 - 57.3%) highlights the need for task-aware calibration and integration of multiple tasks and modalities. Overall, this work establishes a rigorous baseline for VFM-based PD screening and provides a roadmap for selecting suitable tasks and architectures in remote neurological monitoring. Code and anonymized structured data are publicly available: https://anonymous.4open.science/r/parkinson\_video\_benchmarking-A2C5

Keywords

Cite

@article{arxiv.2602.13507,
  title  = {Benchmarking Video Foundation Models for Remote Parkinson's Disease Screening},
  author = {Md Saiful Islam and Ekram Hossain and Abdelrahman Abdelkader and Tariq Adnan and Fazla Rabbi Mashrur and Sooyong Park and Praveen Kumar and Qasim Sudais and Natalia Chunga and Nami Shah and Jan Freyberg and Christopher Kanan and Ruth Schneider and Ehsan Hoque},
  journal= {arXiv preprint arXiv:2602.13507},
  year   = {2026}
}
R2 v1 2026-07-01T10:36:22.528Z