Batch inference for agentic workloads stresses the GPU key-value (KV) cache in a sustained and cumulative manner, often causing severe throughput degradation well before memory capacity is exhausted. We identify this phenomenon as middle-phase thrashing, a previously under-characterized pathology in which cache efficiency collapses as long-lived agents accumulate state over time. We argue that mitigating this pathology requires moving beyond reactive, request-level cache management to proactive, agent-level admission control. Drawing inspiration from congestion control in distributed systems, we view the KV cache as a shared resource whose efficient utilization depends on feedback-driven regulation. Based on this insight, we present CONCUR, a lightweight control layer that regulates agent admission to bound aggregate cache pressure while preserving execution continuity. CONCUR adapts a cache-aware control algorithm to dynamically adjust the number of active agents using runtime cache signals. Across large models and real-world agent workloads, CONCUR prevents middle-phase thrashing and improves batch inference throughput by up to 4.09x on Qwen3-32B and 1.9x on DeepSeek-V3, while remaining compatible with existing LLM serving systems.
@article{arxiv.2601.22705,
title = {CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control},
author = {Qiaoling Chen and Zhisheng Ye and Tian Tang and Peng Sun and Boyu Tian and Guoteng Wang and Shenggui Li and Yonggang Wen and Zhenhua Han and Tianwei Zhang},
journal= {arXiv preprint arXiv:2601.22705},
year = {2026}
}