English

CL-bench Life: Can Language Models Learn from Real-Life Context?

Computation and Language 2026-05-01 v1

Abstract

Today's AI assistants such as OpenClaw are designed to handle context effectively, making context learning an increasingly important capability for models. As these systems move beyond professional settings into everyday life, the nature of the contexts they must handle also shifts. Real-life contexts are often messy, fragmented, and deeply tied to personal and social experience, such as multi-party conversations, personal archives, and behavioral traces. Yet it remains unclear whether current frontier language models can reliably learn from such contexts and solve tasks grounded in them. To this end, we introduce CL-bench Life, a fully human-curated benchmark comprising 405 context-task pairs and 5,348 verification rubrics, covering common real-life scenarios. Solving tasks in CL-bench Life requires models to reason over complex, messy real-life contexts, calling for strong real-life context learning abilities that go far beyond those evaluated in existing benchmarks. We evaluate ten frontier LMs and find that real-life context learning remains highly challenging: even the best-performing model achieves only 19.3% task solving rate, while the average performance across models is only 13.8%. Models still struggle to reason over contexts such as messy group chat histories and fragmented behavioral records from everyday life. CL-bench Life provides a crucial testbed for advancing real-life context learning, and progress on it can enable more intelligent and reliable AI assistants in everyday life.

Keywords

Cite

@article{arxiv.2604.27043,
  title  = {CL-bench Life: Can Language Models Learn from Real-Life Context?},
  author = {Shihan Dou and Yujiong Shen and Chenhao Huang and Junjie Ye and Jiayi Chen and Junzhe Wang and Qianyu He and Shichun Liu and Changze Lv and Jiahang Lin and Jiazheng Zhang and Ming Zhang and Shaofan Liu and Tao Ji and Zhangyue Yin and Cheng Zhang and Huaibing Xie and Jianglu Hu and Jingcheng Deng and Lincheng Li and Minda Hu and Shaolei Wang and Syrus Zhao and Weichao Wang and Yan Lei and Yang Liu and Yanling Xiao and Yiting Liu and Zenan Xu and Zhen Guo and Ziliang Zhao and Pluto Zhou and Tao Gui and Qi Zhang and Xuanjing Huang and Yu-Gang Jiang and Di Wang and Shunyu Yao},
  journal= {arXiv preprint arXiv:2604.27043},
  year   = {2026}
}

Comments

50 pages, 11 figures

R2 v1 2026-07-01T12:42:07.159Z