In commercial web search, aligning content freshness with user intent remains challenging due to the highly varied lifespans of information. Traditional industrial approaches rely on static time-window filtering, resulting in "one-size-fits-all" rankings where content may be chronologically recent but semantically expired. To address the limitation, we present a novel Large Language Models (LLMs)-based Query-Aware Dynamic Content Expiration Prediction Framework deployed in Baidu search, reformulating timeliness as a dynamic validity inference task. Our framework extracts fine-grained temporal contexts from documents and leverages LLMs to deduce a query-specific "validity horizon"-a semantic boundary defining when information becomes obsolete based on user intent. Integrated with robust hallucination mitigation strategies to ensure reliability, our approach has been evaluated through offline and online A/B testing on live production traffic. Results demonstrate significant improvements in search freshness and user experience metrics, validating the effectiveness of LLM-driven reasoning for solving semantic expiration at an industrial scale.
@article{arxiv.2605.13052,
title = {RAG-Enhanced Large Language Models for Dynamic Content Expiration Prediction in Web Search},
author = {Tingyu Chen and Wenkai Zhang and Li Gao and Lixin Su and Ge Chen and Dawei Yin and Daiting Shi},
journal= {arXiv preprint arXiv:2605.13052},
year = {2026}
}
Comments
Accepted at SIGIR 2026. Final version: https://doi.org/10.1145/3805712.3808457