English

KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions

Artificial Intelligence 2026-04-21 v2 Information Retrieval

Abstract

Existing long-horizon memory benchmarks mostly use multi-turn dialogues or synthetic user histories, which makes retrieval performance an imperfect proxy for person understanding. We present \BenchName, a publicly releasable benchmark built from long-form autobiographical narratives, where actions, context, and inner thoughts provide dense evidence for inferring stable motivations and decision principles. \BenchName~reconstructs each narrative into a flashback-aware, time-anchored stream and evaluates models with evidence-linked questions spanning factual recall, subjective state attribution, and principle-level reasoning. Across diverse narrative sources, retrieval-augmented systems mainly improve factual accuracy, while errors persist on temporally grounded explanations and higher-level inferences, highlighting the need for memory mechanisms beyond retrieval. Our data is in \href{KnowMeBench}{https://github.com/QuantaAlpha/KnowMeBench}.

Keywords

Cite

@article{arxiv.2601.04745,
  title  = {KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions},
  author = {Tingyu Wu and Zhisheng Chen and Ziyan Weng and Shuhe Wang and Chenglong Li and Shuo Zhang and Sen Hu and Silin Wu and Qizhen Lan and Huacan Wang and Ronghao Chen},
  journal= {arXiv preprint arXiv:2601.04745},
  year   = {2026}
}