English

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

Computer Vision and Pattern Recognition 2025-08-20 v1

Abstract

The aspiration for artificial general intelligence, fueled by the rapid progress of multimodal models, demands human-comparable performance across diverse environments. We propose HumanPCR, an evaluation suite for probing MLLMs' capacity about human-related visual contexts across three hierarchical levels: Perception, Comprehension, and Reasoning (denoted by Human-P, Human-C, and Human-R, respectively). Human-P and Human-C feature over 6,000 human-verified multiple choice questions, assessing massive tasks of 9 dimensions, including but not limited to essential skills frequently overlooked by existing benchmarks. Human-R offers a challenging manually curated video reasoning test that requires integrating multiple visual evidences, proactively extracting context beyond question cues, and applying human-like expertise. Each question includes human-annotated Chain-of-Thought (CoT) rationales with key visual evidence to support further research. Extensive evaluations on over 30 state-of-the-art models exhibit significant challenges in human-centric visual understanding, particularly in tasks involving detailed space perception, temporal understanding, and mind modeling. Moreover, analysis of Human-R reveals the struggle of models in extracting essential proactive visual evidence from diverse human scenes and their faulty reliance on query-guided retrieval. Even with advanced techniques like scaling visual contexts and test-time thinking yield only limited benefits. We hope HumanPCR and our findings will advance the development, evaluation, and human-centric application of multimodal models.

Keywords

Cite

@article{arxiv.2508.13692,
  title  = {HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes},
  author = {Keliang Li and Hongze Shen and Hao Shi and Ruibing Hou and Hong Chang and Jie Huang and Chenghao Jia and Wen Wang and Yiling Wu and Dongmei Jiang and Shiguang Shan and Xilin Chen},
  journal= {arXiv preprint arXiv:2508.13692},
  year   = {2025}
}
R2 v1 2026-07-01T04:56:29.158Z