Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations
Abstract
Large language models (LLMs) are susceptible to a variety of risks, from non-faithful output to biased and toxic generations. Due to several limiting factors surrounding LLMs (training cost, API access, data availability, etc.), it may not always be feasible to impose direct safety constraints on a deployed model. Therefore, an efficient and reliable alternative is required. To this end, we present our ongoing efforts to create and deploy a library of detectors: compact and easy-to-build classification models that provide labels for various harms. In addition to the detectors themselves, we discuss a wide range of uses for these detector models - from acting as guardrails to enabling effective AI governance. We also deep dive into inherent challenges in their development and discuss future work aimed at making the detectors more reliable and broadening their scope.
Cite
@article{arxiv.2403.06009,
title = {Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations},
author = {Swapnaja Achintalwar and Adriana Alvarado Garcia and Ateret Anaby-Tavor and Ioana Baldini and Sara E. Berger and Bishwaranjan Bhattacharjee and Djallel Bouneffouf and Subhajit Chaudhury and Pin-Yu Chen and Lamogha Chiazor and Elizabeth M. Daly and Kirushikesh DB and Rogério Abreu de Paula and Pierre Dognin and Eitan Farchi and Soumya Ghosh and Michael Hind and Raya Horesh and George Kour and Ja Young Lee and Nishtha Madaan and Sameep Mehta and Erik Miehling and Keerthiram Murugesan and Manish Nagireddy and Inkit Padhi and David Piorkowski and Ambrish Rawat and Orna Raz and Prasanna Sattigeri and Hendrik Strobelt and Sarathkrishna Swaminathan and Christoph Tillmann and Aashka Trivedi and Kush R. Varshney and Dennis Wei and Shalisha Witherspooon and Marcel Zalmanovici},
journal= {arXiv preprint arXiv:2403.06009},
year = {2024}
}