English

Addressable Recall Compaction for Long Context-Window Control in AI Agents

Artificial Intelligence 2026-07-27 v1 Computation and Language

Abstract

Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window. Existing compaction methods address this limitation by discarding, summarizing, or retrieving earlier information, but they may remove task-critical details or fail to recover them reliably. We propose ARC (Addressable Recall Compaction), a context-management framework that separates archival storage from active-context presentation. ARC stores tool observations in an append-only, ID-addressable log and replaces older observations with compact citations when compaction is required. The agent can subsequently use these identifiers to request stored content without re-executing the corresponding tools or depending solely on similarity-based retrieval. We evaluate ARC using Qwen3-8B with a 16k context window and Qwen3-32B with a 32k context window. On the Needle-in-a-Haystack evaluation, ARC achieves an average exact-answer accuracy of 99.40%, compared with 88.12% for the best-performing baseline in our evaluation. ARC also reduces estimated serving time and HBM traffic under our hardware-cost model. On the LongBench-v2 Hard subset, ARC obtains an average accuracy of 29.97%, compared with 28.25% for the best-performing baseline. These results indicate that explicit, address-based recall can improve information retention and serving efficiency relative to the evaluated context-management baselines under the tested settings.

Cite

@article{arxiv.2607.25066,
  title  = {Addressable Recall Compaction for Long Context-Window Control in AI Agents},
  author = {Thang Dang and Yuma Ichikawa and Sakina Fatima and Koichi Shirahata},
  journal= {arXiv preprint arXiv:2607.25066},
  year   = {2026}
}

Comments

20 pages, 2 figures