English

Shape vs. Context: Examining Human--AI Gaps in Ambiguous Japanese Character Recognition

Human-Computer Interaction 2026-03-02 v1 Computer Vision and Pattern Recognition

Abstract

High text recognition performance does not guarantee that Vision-Language Models (VLMs) share human-like decision patterns when resolving ambiguity. We investigate this behavioral gap by directly comparing humans and VLMs using continuously interpolated Japanese character shapes generated via a β\beta-VAE. We estimate decision boundaries in a single-character recognition (shape-only task) and evaluate whether VLM responses align with human judgments under shape in context (i.e., embedding an ambiguous character near the human decision boundary in word-level context). We find that human and VLM decision boundaries differ in the shape-only task, and that shape in context can improve human alignment in some conditions. These results highlight qualitative behavioral differences, offering foundational insights toward human--VLM alignment benchmarking.

Keywords

Cite

@article{arxiv.2602.23746,
  title  = {Shape vs. Context: Examining Human--AI Gaps in Ambiguous Japanese Character Recognition},
  author = {Daichi Haraguchi},
  journal= {arXiv preprint arXiv:2602.23746},
  year   = {2026}
}

Comments

Accepted to CHI 2026 Poster track

R2 v1 2026-07-01T10:55:06.073Z