English

MawForge: Memory-Bounded Expert Materialization for Local Mixture-of-Experts Inference

Machine Learning 2026-06-17 v1

Abstract

Sparse Mixture-of-Experts (MoE) language models separate total parameter count from per-token active computation, but local inference systems often still require the full model, key-value cache, runtime buffers, and operatingsystem headroom to fit in fast memory. MawForge tests a different systems hypothesis: local MoE serving can be made practical on constrained unified-memory machines by storing the full model on disk, keeping common tensors resident, and materializing routed expert tensors into a bounded execution cache on demand. The central finding is that MawForge is effective as a bounded execution mechanism and measurement substrate for local MoE inference, but not as a cache-maximization policy. Performance depends on balancing expert reuse against resident footprint, KV-cache size, quantization, route locality, and macOS memory pressure.

Cite

@article{arxiv.2607.09686,
  title  = {MawForge: Memory-Bounded Expert Materialization for Local Mixture-of-Experts Inference},
  author = {Craig Opie},
  journal= {arXiv preprint arXiv:2607.09686},
  year   = {2026}
}

Comments

7 pages