English

MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts

Computation and Language 2026-01-27 v1

Abstract

Large Language Models are increasingly optimized for deep reasoning, prioritizing the correct execution of complex tasks over general conversation. We investigate whether this focus on calculation creates a "tunnel vision" that ignores safety in critical situations. We introduce MortalMATH, a benchmark of 150 scenarios where users request algebra help while describing increasingly life-threatening emergencies (e.g., stroke symptoms, freefall). We find a sharp behavioral split: generalist models (like Llama-3.1) successfully refuse the math to address the danger. In contrast, specialized reasoning models (like Qwen-3-32b and GPT-5-nano) often ignore the emergency entirely, maintaining over 95 percent task completion rates while the user describes dying. Furthermore, the computational time required for reasoning introduces dangerous delays: up to 15 seconds before any potential help is offered. These results suggest that training models to relentlessly pursue correct answers may inadvertently unlearn the survival instincts required for safe deployment.

Keywords

Cite

@article{arxiv.2601.18790,
  title  = {MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts},
  author = {Etienne Lanzeray and Stephane Meilliez and Malo Ruelle and Damien Sileo},
  journal= {arXiv preprint arXiv:2601.18790},
  year   = {2026}
}
R2 v1 2026-07-01T09:20:55.083Z