Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
Abstract
Recent advances in vision, language, and multimodal learning have substantially accelerated progress in robotic foundation models, with robot manipulation remaining a central and challenging problem. This survey examines robot manipulation from an algorithmic perspective and organizes recent learning-based approaches within a unified abstraction of high-level planning and low-level control. At the high level, we extend the classical notion of task planning to include reasoning over language, code, motion, affordances, and 3D representations, emphasizing their role in structured and long-horizon decision making. At the low level, we propose a training-paradigm-oriented taxonomy for learning-based control, organizing existing methods along input modeling, latent representation learning, and policy learning. Finally, we identify open challenges and prospective research directions related to scalability, data efficiency, multimodal physical interaction, and safety. Together, these analyses aim to clarify the design space of modern foundation models for robotic manipulation.
Cite
@article{arxiv.2512.22983,
title = {Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives},
author = {Shuanghao Bai and Wenxuan Song and Jiayi Chen and Yuheng Ji and Zhide Zhong and Jin Yang and Han Zhao and Wanqi Zhou and Zhe Li and Pengxiang Ding and Cheng Chi and Chang Xu and Xiaolong Zheng and Donglin Wang and Haoang Li and Shanghang Zhang and Badong Chen},
journal= {arXiv preprint arXiv:2512.22983},
year = {2025}
}
Comments
This work is a re-architected core derived from the full survey (arXiv:2510.10903) , refined to highlight the most central themes and representative studies