English

Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation

Computation and Language 2024-10-01 v1

Abstract

Background: Large language models (LLMs) are trained to follow directions, but this introduces a vulnerability to blindly comply with user requests even if they generate wrong information. In medicine, this could accelerate the generation of misinformation that impacts human well-being. Objectives/Methods: We analyzed compliance to requests to generate misleading content about medications in settings where models know the request is illogical. We investigated whether in-context directions and instruction-tuning of LLMs to prioritize logical reasoning over compliance reduced misinformation risk. Results: While all frontier LLMs complied with misinformation requests, both prompt-based and parameter-based approaches can improve the detection of logic flaws in requests and prevent the dissemination of medical misinformation. Conclusion: Shifting LLMs to prioritize logic over compliance could reduce risks of exploitation for medical misinformation.

Keywords

Cite

@article{arxiv.2409.20385,
  title  = {Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation},
  author = {Shan Chen and Mingye Gao and Kuleen Sasse and Thomas Hartvigsen and Brian Anthony and Lizhou Fan and Hugo Aerts and Jack Gallifant and Danielle Bitterman},
  journal= {arXiv preprint arXiv:2409.20385},
  year   = {2024}
}

Comments

Submitted for Review