English

Frame Sampling Strategies Matter: A Benchmark for small vision language models

Computer Vision and Pattern Recognition 2026-03-17 v2 Computation and Language

Abstract

Comparing vision language models on videos is particularly complex, as the performances is jointly determined by the model's visual representation capacity and the frame-sampling strategy used to construct the input. Current video benchmarks are suspected to suffer from substantial frame-sampling bias, as models are evaluated with different frame selection strategies. In this work, we propose the first frame-accurate benchmark of state-of-the-art small VLMs for video question-answering, evaluated under controlled frame-sampling strategies. Our results confirm the suspected bias and highlight both data-specific and task-specific behaviors of SVLMs under different frame-sampling techniques. By open-sourcing our benchmarking code, we provide the community with a reproducible and unbiased protocol for evaluating video VLMs and emphasize the need for standardized frame-sampling strategies tailored to each benchmarking dataset in future research.

Keywords

Cite

@article{arxiv.2509.14769,
  title  = {Frame Sampling Strategies Matter: A Benchmark for small vision language models},
  author = {Marija Brkic and Anas Filali Razzouki and Yannis Tevissen and Khalil Guetari and Mounim A. El Yacoubi},
  journal= {arXiv preprint arXiv:2509.14769},
  year   = {2026}
}
R2 v1 2026-07-01T05:43:27.581Z