English

TimeSeek: Temporal Reliability of Agentic Forecasters

Artificial Intelligence 2026-04-07 v1

Abstract

We introduce TimeSeek, a benchmark for studying how the reliability of agentic LLM forecasters changes over a prediction market's lifecycle. We evaluate 10 frontier models on 150 CFTC-regulated Kalshi binary markets at five temporal checkpoints, with and without web search, for 15,000 forecasts total. Models are most competitive early in a market's life and on high-uncertainty markets, but much less competitive near resolution and on strong-consensus markets. Web search improves pooled Brier Skill Score (BSS) for every model overall, yet hurts in 12% of model-checkpoint pairs, indicating that retrieval is helpful on average but not uniformly so. Simple two-model ensembles reduce error without surpassing the market overall. These descriptive results motivate time-aware evaluation and selective-deference policies rather than a single market snapshot or a uniform tool-use setting.

Keywords

Cite

@article{arxiv.2604.04220,
  title  = {TimeSeek: Temporal Reliability of Agentic Forecasters},
  author = {Hamza Mostafa and Om Shastri and Dennis Lee},
  journal= {arXiv preprint arXiv:2604.04220},
  year   = {2026}
}

Comments

Workshop paper. 11 pages including references

R2 v1 2026-07-01T11:54:38.326Z