Office Comprehension Benchmark
Abstract
We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.docx, .xlsx, .pptx) and their variants. OCB consists of two tracks. File Fidelity Q&A tests structural and visual perception of office artifacts - tables, charts, embedded images, formulas, and app-specific elements such as headers, speaker notes, and named ranges. Domain Q&A tests expert-level reasoning grounded in real-world industry documents across 12 professional domains, with queries requiring multi-step analysis and synthesis across documents. Each reference answer is decomposed into atomic, binary-gradable claims, and an ensemble of LLM judges scores responses against each claim independently. Even the strongest frontier system in its default reasoning mode reaches only about 59.3% on Domain Q&A; increasing thinking depth within a tier does not move performance materially, while moving to a higher product tier yields modest gains. We release the dataset, evaluation tooling, judge prompt, and a public leaderboard.
Cite
@article{arxiv.2607.01245,
title = {Office Comprehension Benchmark},
author = {Firoz Shaik and Mateus Picanço Lima Gomes and Tanvir Aumi and Jingci Wang and Milos Milunovic and Filip Basara and Ivana Jovanovic and Vishwas Suryanarayanan and Neha Nandan Kenkare and Weiyao Xie and Zhipeng Han and Zheng Zhang and Waleed Shahid and Jay Rathi and Russell Scherer and Thong Q. Nguyen and Michael Bentley and Tamara Stankovic and Rasika Chakravarthy and Vishal Chowdhary},
journal= {arXiv preprint arXiv:2607.01245},
year = {2026}
}