English

Efficient Dataset Selection for Continual Adaptation of Generative Recommenders

Information Retrieval 2026-04-10 v1 Machine Learning

Abstract

Recommendation systems must continuously adapt to evolving user behavior, yet the volume of data generated in large-scale streaming environments makes frequent full retraining impractical. This work investigates how targeted data selection can mitigate performance degradation caused by temporal distributional drift while maintaining scalability. We evaluate a range of representation choices and sampling strategies for curating small but informative subsets of user interaction data. Our results demonstrate that gradient-based representations, coupled with distribution-matching, improve downstream model performance, achieving training efficiency gains while preserving robustness to drift. These findings highlight data curation as a practical mechanism for scalable monitoring and adaptive model updates in production-scale recommendation systems.

Keywords

Cite

@article{arxiv.2604.07739,
  title  = {Efficient Dataset Selection for Continual Adaptation of Generative Recommenders},
  author = {Cathy Jiao and Juan Elenter and Praveen Ravichandran and Bernd Huber and Joseph Cauteruccio and Todd Wasson and Timothy Heath and Chenyan Xiong and Mounia Lalmas and Paul Bennett},
  journal= {arXiv preprint arXiv:2604.07739},
  year   = {2026}
}

Comments

ICLR 2026 CAO Workshop (Oral)

R2 v1 2026-07-01T12:00:26.269Z