English

A Survey of Generative Information Retrieval

Information Retrieval 2024-06-05 v2 Computation and Language

Abstract

Generative Retrieval (GR) is an emerging paradigm in information retrieval that leverages generative models to directly map queries to relevant document identifiers (DocIDs) without the need for traditional query processing or document reranking. This survey provides a comprehensive overview of GR, highlighting key developments, indexing and retrieval strategies, and challenges. We discuss various document identifier strategies, including numerical and string-based identifiers, and explore different document representation methods. Our primary contribution lies in outlining future research directions that could profoundly impact the field: improving the quality of query generation, exploring learnable document identifiers, enhancing scalability, and integrating GR with multi-task learning frameworks. By examining state-of-the-art GR techniques and their applications, this survey aims to provide a foundational understanding of GR and inspire further innovations in this transformative approach to information retrieval. We also make the complementary materials such as paper collection publicly available at https://github.com/MiuLab/GenIR-Survey/

Keywords

Cite

@article{arxiv.2406.01197,
  title  = {A Survey of Generative Information Retrieval},
  author = {Tzu-Lin Kuo and Tzu-Wei Chiu and Tzung-Sheng Lin and Sheng-Yang Wu and Chao-Wei Huang and Yun-Nung Chen},
  journal= {arXiv preprint arXiv:2406.01197},
  year   = {2024}
}
R2 v1 2026-06-28T16:50:54.832Z