English

What's the next frontier for Data-centric AI? Data Savvy Agents

Machine Learning 2025-11-04 v1

Abstract

The recent surge in AI agents that autonomously communicate, collaborate with humans and use diverse tools has unlocked promising opportunities in various real-world settings. However, a vital aspect remains underexplored: how agents handle data. Scalable autonomy demands agents that continuously acquire, process, and evolve their data. In this paper, we argue that data-savvy capabilities should be a top priority in the design of agentic systems to ensure reliable real-world deployment. Specifically, we propose four key capabilities to realize this vision: (1) Proactive data acquisition: enabling agents to autonomously gather task-critical knowledge or solicit human input to address data gaps; (2) Sophisticated data processing: requiring context-aware and flexible handling of diverse data challenges and inputs; (3) Interactive test data synthesis: shifting from static benchmarks to dynamically generated interactive test data for agent evaluation; and (4) Continual adaptation: empowering agents to iteratively refine their data and background knowledge to adapt to shifting environments. While current agent research predominantly emphasizes reasoning, we hope to inspire a reflection on the role of data-savvy agents as the next frontier in data-centric AI.

Keywords

Cite

@article{arxiv.2511.01015,
  title  = {What's the next frontier for Data-centric AI? Data Savvy Agents},
  author = {Nabeel Seedat and Jiashuo Liu and Mihaela van der Schaar},
  journal= {arXiv preprint arXiv:2511.01015},
  year   = {2025}
}

Comments

Presented at ICLR 2025 Data-FM. Seedat & Liu contributed equally

R2 v1 2026-07-01T07:18:12.962Z