${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
Abstract
We present a new robotic foundation model, called , that can enable strong out-of-the-box performance in a wide range of scenarios. can follow diverse language instructions in unseen environments, including multi-stage tasks with various kitchen appliances, provide zero-shot cross-embodiment generalization, for example enabling a robot to fold laundry without seeing the task before, and perform challenging tasks such as operating an espresso machine out of the box at a level of performance that matches much more specialized RL-finetuned models. The main idea behind is to use diverse context conditioning during training. This conditioning information, contained in the prompt, makes it possible to steer the model precisely to perform many tasks with different strategies. It is conditioned not just on a language command that describes what it should do, but on additional multimodal information that also describes the manner or strategy in which it should do it, including metadata about task performance and subgoal images. This enables to use very diverse data, including demonstrations, potentially suboptimal (autonomous) data including failures, and data from non-robot sources. Our experiments evaluate across numerous tasks with multiple robot platforms, on tasks that require speed and dexterity, language following, and compositional task generalization.
Cite
@article{arxiv.2604.15483,
title = {${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities},
author = {Physical Intelligence and Bo Ai and Ali Amin and Raichelle Aniceto and Ashwin Balakrishna and Greg Balke and Kevin Black and George Bokinsky and Shihao Cao and Thomas Charbonnier and Vedant Choudhary and Foster Collins and Ken Conley and Grace Connors and James Darpinian and Karan Dhabalia and Maitrayee Dhaka and Jared DiCarlo and Danny Driess and Michael Equi and Adnan Esmail and Yunhao Fang and Chelsea Finn and Catherine Glossop and Thomas Godden and Ivan Goryachev and Lachlan Groom and Haroun Habeeb and Hunter Hancock and Karol Hausman and Gashon Hussein and Victor Hwang and Brian Ichter and Connor Jacobsen and Szymon Jakubczak and Rowan Jen and Tim Jones and Gregg Kammerer and Ben Katz and Liyiming Ke and Mairbek Khadikov and Chandra Kuchi and Marinda Lamb and Devin LeBlanc and Brendon LeCount and Sergey Levine and Xinyu Li and Adrian Li-Bell and Vladislav Lialin and Zhonglin Liang and Wallace Lim and Yao Lu and Enyu Luo and Vishnu Mano and Nandan Marwaha and Aikys Mongush and Liam Murphy and Suraj Nair and Tyler Patterson and Karl Pertsch and Allen Z. Ren and Gavin Schelske and Charvi Sharma and Baifeng Shi and Lucy Xiaoyang Shi and Laura Smith and Jost Tobias Springenberg and Kyle Stachowicz and Will Stoeckle and Jiaming Tang and Jimmy Tanner and Shalom Tekeste and Marcel Torne and Kyle Vedder and Quan Vuong and Anna Walling and Haohuan Wang and Jason Wang and XuDong Wang and Chris Whalen and Samuel Whitmore and Blake Williams and Charles Xu and Sukwon Yoo and Lili Yu and Wuming Zhang and Zhuoyang Zhang and Ury Zhilinsky},
journal= {arXiv preprint arXiv:2604.15483},
year = {2026}
}
Comments
Website: https://www.pi.website/blog/pi07