English
Related papers

Related papers: Proc4Gem: Foundation models for physical agency th…

200 papers

This paper presents a system for procedurally generating agent-based narratives using large language models (LLMs). Users could drag and drop multiple agents and objects into a scene, with each entity automatically assigned semantic…

Graphics · Computer Science 2025-12-24 Vinayak Regmi , Christos Mousas

Embodied agents, in the form of virtual agents or social robots, are rapidly becoming more widespread. In human-human interactions, humans use nonverbal behaviours to convey their attitudes, feelings, and intentions. Therefore, this…

Artificial Intelligence · Computer Science 2026-04-30 Carson Yu Liu , Gelareh Mohammadi , Yang Song , Wafa Johal

The rise of embodied AI has greatly improved the possibility of general mobile agent systems. At present, many evaluation platforms with rich scenes, high visual fidelity and various application scenarios have been developed. In this paper,…

Robotics · Computer Science 2024-10-30 Haoran Li , Shasha Liu , Mingjun Ma , Guangzheng Hu , Yaran Chen , Dongbin Zhao

Desktop organization remains challenging for service robots because of heterogeneous objects and diverse manipulation objectives, such as collection and stacking. In this article, a task-oriented framework is presented for organizing planar…

Robotics · Computer Science 2026-05-12 Yi Dong , Yangjun Liu , Jinjun Duan , Yang Li , Zhendong Dai

We address the challenging problem of robotic grasping and manipulation in the presence of uncertainty. This uncertainty is due to noisy sensing, inaccurate models and hard-to-predict environment dynamics. We quantify the importance of…

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent…

In model-based reinforcement learning, generative and temporal models of environments can be leveraged to boost agent performance, either by tuning the agent's representations during training or via use as part of an explicit planning…

Recently developed pretrained models can encode rich world knowledge expressed in multiple modalities, such as text and images. However, the outputs of these models cannot be integrated into algorithms to solve sequential decision-making…

Artificial Intelligence · Computer Science 2024-06-19 Yunhao Yang , Cyrus Neary , Ufuk Topcu

Despite a widespread success in various applications, large language models (LLMs) often stumble when tackling basic physical reasoning or executing robotics tasks, due to a lack of direct experience with the physical nuances of the real…

Computation and Language · Computer Science 2024-11-13 Haolan Liu , Jishen Zhao

Humanoid robots are expected to execute agile and expressive whole-body motions in real-world settings. Existing text-to-motion generation models are predominantly trained on captured human motion datasets, whose priors assume human…

Evaluating the surroundings to gain understanding, frame perspectives, and anticipate behavioral reactions is an inherent human trait. However, these continuous encounters are diverse and complex, posing challenges to their study and…

Computers and Society · Computer Science 2026-02-26 Deepank Verma , Olaf Mumm , Vanessa Miriam Carlow

Video generation models are rapidly improving in their ability to synthesize human actions in novel contexts, holding the potential to serve as high-level planners for contextual robot control. To realize this potential, a key research…

Robotics · Computer Science 2025-12-12 James Ni , Zekai Wang , Wei Lin , Amir Bar , Yann LeCun , Trevor Darrell , Jitendra Malik , Roei Herzig

Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. However, scaling such agent training is limited by the lack of diverse and reliable environments.…

Artificial Intelligence · Computer Science 2026-05-26 Zhaoyang Wang , Canwen Xu , Boyi Liu , Yite Wang , Siwei Han , Zhewei Yao , Huaxiu Yao , Yuxiong He

We present a learnable physics-based predictive model that provides accurate motion and force-torque prediction of the robot end effector in contact-rich manipulation. The proposed model extends the state-of-the-art GNN-based simulator…

Robotics · Computer Science 2026-03-03 Zongyao Yi , Joachim Hertzberg , Martin Atzmueller

Sim-to-real transfer remains a critical bottleneck for deploying dexterous manipulation policies learned in simulation to real-world robots. Existing approaches rely on manually designed domain randomization or task-specific adaptation,…

Robotics · Computer Science 2026-05-08 Zijian Zeng , Fei Ding , Huiming Yang , Xianwei Li , Yuhao Liao

Pretrained Foundation Models (PFMs) are regarded as the foundation for various downstream tasks with different data modalities. A PFM (e.g., BERT, ChatGPT, and GPT-4) is trained on large-scale data which provides a reasonable parameter…

Realistic fine-grained multi-agent simulation of real-world complex systems is crucial for many downstream tasks such as reinforcement learning. Recent work has used generative models (GANs in particular) for providing high-fidelity…

Machine Learning · Computer Science 2022-02-25 Changyu Chen , Avinandan Bose , Shih-Fen Cheng , Arunesh Sinha

Differentiable simulators provide analytic gradients, enabling more sample-efficient learning algorithms and paving the way for data intensive learning tasks such as learning from images. In this work, we demonstrate that locomotion…

Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as robots remains a significant challenge. This report…

Robotics · Computer Science 2025-03-27 Gemini Robotics Team , Saminda Abeyruwan , Joshua Ainslie , Jean-Baptiste Alayrac , Montserrat Gonzalez Arenas , Travis Armstrong , Ashwin Balakrishna , Robert Baruch , Maria Bauza , Michiel Blokzijl , Steven Bohez , Konstantinos Bousmalis , Anthony Brohan , Thomas Buschmann , Arunkumar Byravan , Serkan Cabi , Ken Caluwaerts , Federico Casarini , Oscar Chang , Jose Enrique Chen , Xi Chen , Hao-Tien Lewis Chiang , Krzysztof Choromanski , David D'Ambrosio , Sudeep Dasari , Todor Davchev , Coline Devin , Norman Di Palo , Tianli Ding , Adil Dostmohamed , Danny Driess , Yilun Du , Debidatta Dwibedi , Michael Elabd , Claudio Fantacci , Cody Fong , Erik Frey , Chuyuan Fu , Marissa Giustina , Keerthana Gopalakrishnan , Laura Graesser , Leonard Hasenclever , Nicolas Heess , Brandon Hernaez , Alexander Herzog , R. Alex Hofer , Jan Humplik , Atil Iscen , Mithun George Jacob , Deepali Jain , Ryan Julian , Dmitry Kalashnikov , M. Emre Karagozler , Stefani Karp , Chase Kew , Jerad Kirkland , Sean Kirmani , Yuheng Kuang , Thomas Lampe , Antoine Laurens , Isabel Leal , Alex X. Lee , Tsang-Wei Edward Lee , Jacky Liang , Yixin Lin , Sharath Maddineni , Anirudha Majumdar , Assaf Hurwitz Michaely , Robert Moreno , Michael Neunert , Francesco Nori , Carolina Parada , Emilio Parisotto , Peter Pastor , Acorn Pooley , Kanishka Rao , Krista Reymann , Dorsa Sadigh , Stefano Saliceti , Pannag Sanketi , Pierre Sermanet , Dhruv Shah , Mohit Sharma , Kathryn Shea , Charles Shu , Vikas Sindhwani , Sumeet Singh , Radu Soricut , Jost Tobias Springenberg , Rachel Sterneck , Razvan Surdulescu , Jie Tan , Jonathan Tompson , Vincent Vanhoucke , Jake Varley , Grace Vesom , Giulia Vezzani , Oriol Vinyals , Ayzaan Wahid , Stefan Welker , Paul Wohlhart , Fei Xia , Ted Xiao , Annie Xie , Jinyu Xie , Peng Xu , Sichun Xu , Ying Xu , Zhuo Xu , Yuxiang Yang , Rui Yao , Sergey Yaroshenko , Wenhao Yu , Wentao Yuan , Jingwei Zhang , Tingnan Zhang , Allan Zhou , Yuxiang Zhou

We explore building generative neural network models of popular reinforcement learning environments. Our world model can be trained quickly in an unsupervised manner to learn a compressed spatial and temporal representation of the…

Machine Learning · Computer Science 2018-05-10 David Ha , Jürgen Schmidhuber