中文

HuM-Eval:面向人类中心视频评估的粗到细框架

计算机视觉与模式识别 2026-04-29 v1

摘要

视频生成模型在最近几年取得了快速发展,其中生成自然人类运动起到了关键作用。然而,准确评估生成人类运动视频的质量仍然是一个 significant 挑战。现有的 evaluation metrics 主要关注 global scene statistics,常常忽略细粒度的人类细节,从而无法与人类主观偏好保持一致。为桥接 this gap,我们提出了 HuM-Eval,一个 novel 人类中心 evaluation 框架,采用 coarse-to-fine 策略。具体而言,我们的框架首先利用 Vision Language Model 对 global video quality 进行粗略评估。随后进行细粒度分析,使用 2D pose 验证解剖正确性,使用 3D human motion 评估 motion stability。大量实验表明,HuM-Eval 在 average human correlation 上达到 58.2%,优于 state-of-the-art baselines。此外,我们引入了 HuM-Bench,一个包含 1,000 个 diverse prompts 的 comprehensive benchmark,对现有的 text-to-video 模型进行了详细评估,为 next-generation human motion generation 铺平了道路。

关键词

引用

@article{arxiv.2604.25361,
  title  = {HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation},
  author = {Bingzi Zhang and Kaisi Guan and Ruihua Song},
  journal= {arXiv preprint arXiv:2604.25361},
  year   = {2026}
}

备注

Accepted to the 2026 IEEE International Conference on Multimedia and Expo (ICME 2026)