English

DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity

Machine Learning 2026-02-13 v1 Artificial Intelligence

Abstract

We present DRACO (Deep Research Accuracy, Completeness, and Objectivity), a benchmark of complex deep research tasks. These tasks, which span 10 domains and draw on information sources from 40 countries, originate from anonymized real-world usage patterns within a large-scale deep research system. Tasks are sampled from a de-identified dataset of Perplexity Deep Research requests, then filtered and augmented to ensure that the tasks are anonymized, open-ended and complex, objectively evaluable, and representative of the broad scope of real-world deep research use cases. Outputs are graded against task-specific rubrics along four dimensions: factual accuracy (accuracy), breadth and depth of analysis (including completeness), presentation quality (including objectivity), and citation quality. DRACO is publicly available at https://hf.co/datasets/perplexity-ai/draco.

Keywords

Cite

@article{arxiv.2602.11685,
  title  = {DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity},
  author = {Joey Zhong and Hao Zhang and Clare Southern and Jeremy Yang and Thomas Wang and Kate Jung and Shu Zhang and Denis Yarats and Johnny Ho and Jerry Ma},
  journal= {arXiv preprint arXiv:2602.11685},
  year   = {2026}
}
R2 v1 2026-07-01T10:33:13.119Z