English

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

Artificial Intelligence 2026-06-10 v1 Machine Learning

Abstract

Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can cause drift in non-updated behavior. Online distillation mitigates this drift by training on model-generated rollouts, yet uniform reference-conditioned distillation provides coarse supervision: it can under-emphasize reference-supported rollout tokens and supervise omitted facts only indirectly. We introduce RoCo-ACE, a rollout-conditioned online distillation objective for knowledge injection. RoCo uses same-rollout reference-free/reference-conditioned likelihood contrast to reallocate additional distillation weight to reference-supported rollout tokens, while ACE adds sparse reference-side anchored correction for authoritative anchors omitted from the rollout without full-answer imitation. Across three knowledge-injection settings, six retention benchmarks, multiple baselines, and multiple base models, RoCo-ACE achieves the best injected-knowledge accuracy among compared methods while keeping evaluated retention close to the base model.

Keywords

Cite

@article{arxiv.2607.24771,
  title  = {RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection},
  author = {Yan Hong and Wei Li and Kedong Xiu and Jun Lan and Shuheng Zhou and Zhongcai Lyu and Huijia Zhu and Weiqiang Wang and Jianfu Zhang},
  journal= {arXiv preprint arXiv:2607.24771},
  year   = {2026}
}