English

Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with constraints

Artificial Intelligence 2026-04-20 v3

Abstract

Improving the reliability of large language models (LLMs) is critical for deploying them in real-world scenarios. In this paper, we propose \textbf{Deliberative Searcher}, the first framework to integrate certainty calibration with retrieval-based search for open-domain question answering. The agent performs multi-step reflection and verification over Wikipedia data and is trained with a reinforcement learning algorithm that optimizes for accuracy under a soft reliability constraint. Empirical results show that proposed method improves alignment between model confidence and correctness, leading to more trustworthy outputs. This paper will be continuously updated.

Keywords

Cite

@article{arxiv.2507.16727,
  title  = {Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with constraints},
  author = {Zhenyun Yin and Shujie Wang and Xuhong Wang and Xingjun Ma and Yinchun Wang},
  journal= {arXiv preprint arXiv:2507.16727},
  year   = {2026}
}

Comments

Accepted by ACL 2026

R2 v1 2026-07-01T04:13:42.186Z