English

A Systematic Analysis of Chunking Strategies for Reliable Question Answering

Computation and Language 2026-01-21 v1 Information Retrieval

Abstract

We study how document chunking choices impact the reliability of Retrieval-Augmented Generation (RAG) systems in industry. While practice often relies on heuristics, our end-to-end evaluation on Natural Questions systematically varies chunking method (token, sentence, semantic, code), chunk size, overlap, and context length. We use a standard industrial setup: SPLADE retrieval and a Mistral-8B generator. We derive actionable lessons for cost-efficient deployment: (i) overlap provides no measurable benefit and increases indexing cost; (ii) sentence chunking is the most cost-effective method, matching semantic chunking up to ~5k tokens; (iii) a "context cliff" reduces quality beyond ~2.5k tokens; and (iv) optimal context depends on the goal (semantic quality peaks at small contexts; exact match at larger ones).

Keywords

Cite

@article{arxiv.2601.14123,
  title  = {A Systematic Analysis of Chunking Strategies for Reliable Question Answering},
  author = {Sofia Bennani and Charles Moslonka},
  journal= {arXiv preprint arXiv:2601.14123},
  year   = {2026}
}

Comments

3 pages, 2 figures, 1 table, pre-print