English

SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection

Computation and Language 2024-08-26 v1 Artificial Intelligence Machine Learning

Abstract

Large language models (LLMs) are highly capable but face latency challenges in real-time applications, such as conducting online hallucination detection. To overcome this issue, we propose a novel framework that leverages a small language model (SLM) classifier for initial detection, followed by a LLM as constrained reasoner to generate detailed explanations for detected hallucinated content. This study optimizes the real-time interpretable hallucination detection by introducing effective prompting techniques that align LLM-generated explanations with SLM decisions. Empirical experiment results demonstrate its effectiveness, thereby enhancing the overall user experience.

Keywords

Cite

@article{arxiv.2408.12748,
  title  = {SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection},
  author = {Mengya Hu and Rui Xu and Deren Lei and Yaxi Li and Mingyu Wang and Emily Ching and Eslam Kamal and Alex Deng},
  journal= {arXiv preprint arXiv:2408.12748},
  year   = {2024}
}

Comments

preprint under review

R2 v1 2026-06-28T18:21:31.493Z