English

PushupBench: Your VLM is not good at counting pushups

Computer Vision and Pattern Recognition 2026-04-28 v1 Artificial Intelligence

Abstract

Large vision-language models (VLMs) can recognize \textit{what} happens in video but fail to count \textit{how many} times. We introduce \textbf{PushupBench}, 446 long-form clips (avg. 36.7s) for evaluating repetition counting. The best frontier model achieves 42.1\% exact accuracy; open-source 4B models score \sim6\%, matching supervised baselines. We show that accuracy alone misleads -- weaker models exploit the modal count rather than reason temporally. Fine-tuning on counting with 1k samples transfers to general video understanding: MVBench (+2.15), PerceptionTest (+1.88), TVBench (+4.54), suggesting counting is a proxy for broader temporal reasoning.PushupBench incorporated in \texttt{lmms-eval} (https://github.com/EvolvingLMMs-Lab/lmms-eval/pull/1262) and hosted on (pushupbench.com/)

Keywords

Cite

@article{arxiv.2604.23407,
  title  = {PushupBench: Your VLM is not good at counting pushups},
  author = {Shengzhi Li and Jiarun Chen and Karun Sharma and Jiaqi Su and Shichao Pei},
  journal= {arXiv preprint arXiv:2604.23407},
  year   = {2026}
}
R2 v1 2026-07-01T12:35:17.944Z