硬币的两面:以大语言模型为评估器的幻觉生成与检测
人工智能
2024-07-15 v1 计算与语言
摘要
大语言模型(LLM)中的幻觉检测对于确保其可靠性至关重要。本文 presenting our participation in the CLEF ELOQUENT HalluciGen shared task,其中目标是开发用于生成和检测幻觉内容的评估器。我们探索了四种大语言模型(Llama 3、Gemma、GPT-3.5 Turbo 和 GPT-4)的能力,以实现上述目标。我们还采用了集成多数投票方法,以整合所有四个模型用于检测任务。结果为我们了解这些大语言模型在处理幻觉生成和检测任务方面的优势与不足提供了宝贵的见解。
引用
@article{arxiv.2407.09152,
title = {The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs},
author = {Anh Thu Maria Bui and Saskia Felizitas Brech and Natalie Hußfeldt and Tobias Jennert and Melanie Ullrich and Timo Breuer and Narjes Nikzad Khasmakhi and Philipp Schaer},
journal= {arXiv preprint arXiv:2407.09152},
year = {2024}
}
备注
Paper accepted at ELOQUENT@CLEF'24