English

Evaluating Search Engines and Large Language Models for Answering Health Questions

Information Retrieval 2025-03-07 v3 Artificial Intelligence

Abstract

Search engines (SEs) have traditionally been primary tools for information seeking, but the new Large Language Models (LLMs) are emerging as powerful alternatives, particularly for question-answering tasks. This study compares the performance of four popular SEs, seven LLMs, and retrieval-augmented (RAG) variants in answering 150 health-related questions from the TREC Health Misinformation (HM) Track. Results reveal SEs correctly answer between 50 and 70% of questions, often hindered by many retrieval results not responding to the health question. LLMs deliver higher accuracy, correctly answering about 80% of questions, though their performance is sensitive to input prompts. RAG methods significantly enhance smaller LLMs' effectiveness, improving accuracy by up to 30% by integrating retrieval evidence.

Keywords

Cite

@article{arxiv.2407.12468,
  title  = {Evaluating Search Engines and Large Language Models for Answering Health Questions},
  author = {Marcos Fernández-Pichel and Juan C. Pichel and David E. Losada},
  journal= {arXiv preprint arXiv:2407.12468},
  year   = {2025}
}