English

AcademiClaw: When Students Set Challenges for AI Agents

Artificial Intelligence 2026-05-05 v1 Computers and Society

Abstract

Benchmarks within the OpenClaw ecosystem have thus far evaluated exclusively assistant-level tasks, leaving the academic-level capabilities of OpenClaw largely unexamined. We introduce AcademiClaw, a bilingual benchmark of 80 complex, long-horizon tasks sourced directly from university students' real academic workflows -- homework, research projects, competitions, and personal projects -- that they found current AI agents unable to solve effectively. Curated from 230 student-submitted candidates through rigorous expert review, the final task set spans 25+ professional domains, ranging from olympiad-level mathematics and linguistics problems to GPU-intensive reinforcement learning and full-stack system debugging, with 16 tasks requiring CUDA GPU execution. Each task executes in an isolated Docker sandbox and is scored on task completion by multi-dimensional rubrics combining six complementary techniques, with an independent five-category safety audit providing additional behavioral analysis. Experiments on six frontier models show that even the best achieves only a 55\% pass rate. Further analysis uncovers sharp capability boundaries across task domains, divergent behavioral strategies among models, and a disconnect between token consumption and output quality, providing fine-grained diagnostic signals beyond what aggregate metrics reveal. We hope that AcademiClaw and its open-sourced data and code can serve as a useful resource for the OpenClaw community, driving progress toward agents that are more capable and versatile across the full breadth of real-world academic demands. All data and code are available at https://github.com/GAIR-NLP/AcademiClaw.

Keywords

Cite

@article{arxiv.2605.02661,
  title  = {AcademiClaw: When Students Set Challenges for AI Agents},
  author = {Junjie Yu and Pengrui Lu and Weiye Si and Hongliang Lu and Jiabao Wu and Kaiwen Tao and Kun Wang and Lingyu Yang and Qiran Zhang and Xiuting Guo and Xuanyu Wang and Yang Wang and Yanjie Wang and Yi Yang and Zijian Hu and Ziyi Yang and Zonghan Zhou and Binghao Qiang and Borui Zhang and Chenning Li and Enchang Zhang and Feifan Chen and Feng Jian and Fengyin Sun and Hao Qiu and Hao Zheng and Haoran Zhu and Hongyu Liu and Jianbin Deng and Jiaxin Song and Jiaying Chi and Jiayou Shi and Jie Fang and Jinghui Zhong and Jingyu Zhou and Jinze Li and Junfeng Yi and Junyan Yu and Junzhi Xue and Ni Song and Pengyi Chen and Qi Chen and Quansheng Li and Rui Tao and Shenghai Gong and Shenhang Lu and Tianqi Shen and Tianxiang Zhu and Tiehan Kang and Tingyu Li and Wendi Wu and Xiao Shen and Xiao Zhou and Xiaotao Zhang and Xinrong Li and Xuankun Yang and Xun Zhang and Yan Li and Ye Lu and Yi Wang and Yibo Zhou and Yichi Zhang and Yihao Sun and Yijun Huang and Yixin Zhu and Yixuan Wu and Yuchen Sun and Yue Wu and Yuheng Sun and Yukun Li and Yutian Tu and Yuxuan Qin and Yuzhuo Wu and Zeyu Li and Zhengyu Lou and Zhenning Ran and Zizhu He and Pengfei Liu},
  journal= {arXiv preprint arXiv:2605.02661},
  year   = {2026}
}
R2 v1 2026-07-01T12:48:38.531Z