English

Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare

Computation and Language 2026-01-27 v1

Abstract

As LLMs are increasingly integrated into clinical workflows, their tendency for sycophancy, prioritizing user agreement over factual accuracy, poses significant risks to patient safety. While existing evaluations often rely on subjective datasets, we introduce a robust framework grounded in medical MCQA with verifiable ground truths. We propose the Adjusted Sycophancy Score, a novel metric that isolates alignment bias by accounting for stochastic model instability, or "confusability". Through an extensive scaling analysis of the Qwen-3 and Llama-3 families, we identify a clear scaling trajectory for resilience. Furthermore, we reveal a counter-intuitive vulnerability in reasoning-optimized "Thinking" models: while they demonstrate high vanilla accuracy, their internal reasoning traces frequently rationalize incorrect user suggestions under authoritative pressure. Our results across frontier models suggest that benchmark performance is not a proxy for clinical reliability, and that simplified reasoning structures may offer superior robustness against expert-driven sycophancy.

Keywords

Cite

@article{arxiv.2601.18334,
  title  = {Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare},
  author = {Clément Christophe and Wadood Mohammed Abdul and Prateek Munjal and Tathagata Raha and Ronnie Rajan and Praveenkumar Kanithi},
  journal= {arXiv preprint arXiv:2601.18334},
  year   = {2026}
}