English

ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers

Computation and Language 2026-07-13 v1

Abstract

Large language models are increasingly used to assist scientific reading, but existing evaluation methods often fail to detect whether answers are supported by verifiable citations. We introduce ResearchQA, a benchmark of 6,211 single-paper question-answer pairs from 494 open-access papers spanning eight domains and four question types: lookup, comprehension, multi-hop, and adversarial. ResearchQA is designed for citation-grounded evaluation: it permits multiple valid supporting passages for a claim and rewards grounded refusal when the source paper does not support an answer. We evaluate eight leading closed- and open-weight models in a citation-grounded chat-with-paper setting using a deterministic citation matcher and an LLM-based rubric evaluator. Citation-based metrics separate systems more clearly than LLM-evaluator scores: section coverage and citation accuracy vary substantially across models, while evaluator scores remain tightly compressed. We further find that open-weight models approach the best closed-model citation accuracy while achieving 3 to 6 times lower per-example latency. We release the benchmark, evaluation harness, and evaluator prompt.

Cite

@article{arxiv.2607.11074,
  title  = {ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers},
  author = {Saba Imran and Debanjum Singh Solanky},
  journal= {arXiv preprint arXiv:2607.11074},
  year   = {2026}
}

Comments

19 pages, 9 figures

R2 v1 2026-07-22T20:36:41.108Z