English

Dense Passage Retrieval for Open-Domain Question Answering

Computation and Language 2020-10-02 v3

Abstract

Open-domain question answering relies on efficient passage retrieval to select candidate contexts, where traditional sparse vector space models, such as TF-IDF or BM25, are the de facto method. In this work, we show that retrieval can be practically implemented using dense representations alone, where embeddings are learned from a small number of questions and passages by a simple dual-encoder framework. When evaluated on a wide range of open-domain QA datasets, our dense retriever outperforms a strong Lucene-BM25 system largely by 9%-19% absolute in terms of top-20 passage retrieval accuracy, and helps our end-to-end QA system establish new state-of-the-art on multiple open-domain QA benchmarks.

Keywords

Cite

@article{arxiv.2004.04906,
  title  = {Dense Passage Retrieval for Open-Domain Question Answering},
  author = {Vladimir Karpukhin and Barlas Oğuz and Sewon Min and Patrick Lewis and Ledell Wu and Sergey Edunov and Danqi Chen and Wen-tau Yih},
  journal= {arXiv preprint arXiv:2004.04906},
  year   = {2020}
}

Comments

EMNLP 2020

R2 v1 2026-06-23T14:46:33.860Z