English

EgoNormia: Benchmarking Physical Social Norm Understanding

Computer Vision and Pattern Recognition 2025-06-13 v5 Artificial Intelligence Computation and Language

Abstract

Human activity is moderated by norms; however, supervision for normative reasoning is sparse, particularly where norms are physically- or socially-grounded. We thus present EGONORMIA ϵ\|\epsilon\|, comprising 1,853 (200 for EGONORMIA-verified) multiple choice questions (MCQs) grounded within egocentric videos of human interactions, enabling the evaluation and improvement of normative reasoning in vision-language models (VLMs). EGONORMIA spans seven norm categories: safety, privacy, proxemics, politeness, cooperation, coordination/proactivity, and communication/legibility. To compile this dataset at scale, we propose a novel pipeline to generate grounded MCQs from raw egocentric video. Our work demonstrates that current state-of-the-art VLMs lack robust grounded norm understanding, scoring a maximum of 54% on EGONORMIA and 65% on EGONORMIA-verified, with performance across norm categories indicating significant risks of safety and privacy when VLMs are used in real-world agents. We additionally explore methods for improving normative understanding, demonstrating that a naive retrieval-based generation (RAG) method using EGONORMIA can enhance normative reasoning in VLMs.

Keywords

Cite

@article{arxiv.2502.20490,
  title  = {EgoNormia: Benchmarking Physical Social Norm Understanding},
  author = {MohammadHossein Rezaei and Yicheng Fu and Phil Cuvin and Caleb Ziems and Yanzhe Zhang and Hao Zhu and Diyi Yang},
  journal= {arXiv preprint arXiv:2502.20490},
  year   = {2025}
}

Comments

V4, fixes to title and formatting

R2 v1 2026-06-28T22:00:49.145Z