English

Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

Computational Engineering, Finance, and Science 2026-07-23 v1 Computation and Language

Abstract

Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and largely decoupled from answer correctness: correct answers often coexist with fabricated structural claims absent from the relevant molecules. Yet this does not make the reasoning trace computationally irrelevant. Attribution analyses suggest a shared scratchpad function expressed in model-specific forms: Chem-R and ether-0 rely on fragmented SMILES drafts, whereas ChemDFM-R emphasizes scaffold, positional, and naming cues. Notably, perturbing Chem-R's SMILES sketches degrades generation, showing that structural drafts can be causally load-bearing even when verbal structural claims are largely inert. Together, these results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad. This finding cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.

Keywords

Cite

@article{arxiv.2607.20935,
  title  = {Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad},
  author = {Jiatong Li and Yuxuan Ren and Weida Wang and Xiaoyong Wei and Yatao Bian},
  journal= {arXiv preprint arXiv:2607.20935},
  year   = {2026}
}

Comments

16 pages, 6 figures