English

BAGEL: Benchmarking Animal Knowledge Expertise in Language Models

Computation and Language 2026-04-20 v1 Artificial Intelligence

Abstract

Large language models have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language models handle specialized animal-related knowledge under a unified closed-book evaluation protocol. We introduce BAGEL, a benchmark for evaluating animal knowledge expertise in language models. BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, using a combination of curated examples and automatically generated closed-book question-answer pairs. The benchmark covers multiple aspects of animal knowledge, including taxonomy, morphology, habitat, behavior, vocalization, geographic distribution, and species interactions. By focusing on closed-book evaluation, BAGEL measures animal-related knowledge of models without external retrieval at inference time. BAGEL further supports fine-grained analysis across source domains, taxonomic groups, and knowledge categories, enabling a more precise characterization of model strengths and systematic failure modes. Our benchmark provides a new testbed for studying domain-specific knowledge generalization in language models and for improving their reliability in biodiversity-related applications.

Keywords

Cite

@article{arxiv.2604.16241,
  title  = {BAGEL: Benchmarking Animal Knowledge Expertise in Language Models},
  author = {Jiacheng Shen and Masato Hagiwara and Milad Alizadeh and Ellen Gilsenan-McMahon and Marius Miron and David Robinson and Emmanuel Chemla and Sara Keen and Gagan Narula and Mathieu Laurière and Matthieu Geist and Olivier Pietquin},
  journal= {arXiv preprint arXiv:2604.16241},
  year   = {2026}
}

Comments

28 pages, 3 figures

R2 v1 2026-07-01T12:14:41.145Z