HuM-Eval:面向人类中心视频评估的粗到细框架
计算机视觉与模式识别
2026-04-29 v1
摘要
视频生成模型在最近几年取得了快速发展,其中生成自然人类运动起到了关键作用。然而,准确评估生成人类运动视频的质量仍然是一个 significant 挑战。现有的 evaluation metrics 主要关注 global scene statistics,常常忽略细粒度的人类细节,从而无法与人类主观偏好保持一致。为桥接 this gap,我们提出了 HuM-Eval,一个 novel 人类中心 evaluation 框架,采用 coarse-to-fine 策略。具体而言,我们的框架首先利用 Vision Language Model 对 global video quality 进行粗略评估。随后进行细粒度分析,使用 2D pose 验证解剖正确性,使用 3D human motion 评估 motion stability。大量实验表明,HuM-Eval 在 average human correlation 上达到 58.2%,优于 state-of-the-art baselines。此外,我们引入了 HuM-Bench,一个包含 1,000 个 diverse prompts 的 comprehensive benchmark,对现有的 text-to-video 模型进行了详细评估,为 next-generation human motion generation 铺平了道路。
引用
@article{arxiv.2604.25361,
title = {HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation},
author = {Bingzi Zhang and Kaisi Guan and Ruihua Song},
journal= {arXiv preprint arXiv:2604.25361},
year = {2026}
}
备注
Accepted to the 2026 IEEE International Conference on Multimedia and Expo (ICME 2026)