When Words Predict Workload
Abstract
Standard distributed \ac{llm} schedulers rely on static token counts or rolling latency averages, making them susceptible to failures on statutorily constrained text. On \ac{epo} claims governed by Article 84 \ac{epc}, linguistic rigidity makes human and machine authorship statistically indistinguishable. Resolving this ambiguity mid-flight forces dynamic multi-model ensemble expansion, triggering unpredictable KV-cache and weight-allocation spikes that saturate consumer-grade edge GPU VRAM and cause severe \ac{oom} crashes. To prevent hardware collapse, we propose a CPU-side Linguistic Resource Forecasting (LRF) gateway. The gateway extracts a 16-dimensional text-structure vector and applies an XGBoost predictor to forecast trap-band membership. The resulting escalation probability () is evaluated against a dynamic, closed-form routing threshold () computed via real-time latency telemetry. Requests are safely routed to either a local Qwen2.5-7B edge worker or a remote contrastive ensemble (Qwen2.5 7B + 32B) on an NVIDIA H100 \emph{before} any edge GPU memory is allocated. In a 6,000-request live trial, the LRF gateway reduced the operational misroute fraction () to --, an order of magnitude below the token-count baseline (). Peak edge VRAM remained safely bounded at (under the ceiling) across a variation in \ac{wan} delay. The predictor achieved a live-trial AUROC of , and the dynamic controller yielded an relative reduction in misroutes compared to an equivalent static threshold.
Cite
@article{arxiv.2607.04951,
title = {When Words Predict Workload},
author = {Anubhab Banerjee},
journal= {arXiv preprint arXiv:2607.04951},
year = {2026}
}
Comments
This work has been submitted to the IEEE for possible publication. Permission from the author must be obtained for all uses