English
Related papers

Related papers: GE-Sim 2.0: A Roadmap Towards Comprehensive Closed…

200 papers

With the rapid development of embodied artificial intelligence, significant progress has been made in vision-language-action (VLA) models for general robot decision-making. However, the majority of existing VLAs fail to account for the…

Robotics · Computer Science 2025-02-17 Hongyin Zhang , Pengxiang Ding , Shangke Lyu , Ying Peng , Donglin Wang

Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive perception for decision making. Progress on this question has…

Robotic manipulation policies are advancing rapidly, but their direct evaluation in the real world remains costly, time-consuming, and difficult to reproduce, particularly for tasks involving deformable objects. Simulation provides a…

World models allow autonomous agents to plan and explore by predicting the visual outcomes of different actions. However, for robot manipulation, it is challenging to accurately model the fine-grained robot-object interaction within the…

Robotics · Computer Science 2025-07-30 Fangqi Zhu , Hongtao Wu , Song Guo , Yuxiao Liu , Chilam Cheang , Tao Kong

In this paper, we explore generalizable, perception-to-action robotic manipulation for precise, contact-rich tasks. In particular, we contribute a framework for closed-loop robotic manipulation that automatically handles a category of…

Robotics · Computer Science 2021-02-15 Wei Gao , Russ Tedrake

Robotic simulation today remains challenging to scale up due to the human efforts required to create diverse simulation tasks and scenes. Simulation-trained policies also face scalability issues as many sim-to-real methods focus on a single…

Robotics · Computer Science 2024-10-07 Pu Hua , Minghuan Liu , Annabella Macaluso , Yunfeng Lin , Weinan Zhang , Huazhe Xu , Lirui Wang

Realizing generalizable dynamic object manipulation on conveyor systems is important for enhancing manufacturing efficiency, as it eliminates specialized engineering for different scenarios. To this end, imitation learning emerges as a…

Robotics · Computer Science 2026-02-10 Zhuoling Li , Jinrong Yang , Yong Zhao , Liangliang Ren , Xiaoyang Wu , Zhenhua Xu , Hengshuang Zhao

Generative models offer a scalable and flexible paradigm for simulating complex environments, yet current approaches fall short in addressing the domain-specific requirements of autonomous driving - such as multi-agent interactions,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Lloyd Russell , Anthony Hu , Lorenzo Bertoni , George Fedoseev , Jamie Shotton , Elahe Arani , Gianluca Corrado

Robotic manipulation with deformable objects represents a data-intensive regime in embodied learning, where shape, contact, and topology co-evolve in ways that far exceed the variability of rigids. Although simulation promises relief from…

Physics-aware driving world model is essential for drive planning, out-of-distribution data synthesis, and closed-loop evaluation. However, existing methods often rely on a single diffusion model to directly map driving actions to videos,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zhenya Yang , Zhe Liu , Yuxiang Lu , Liping Hou , Chenxuan Miao , Siyi Peng , Bailan Feng , Xiang Bai , Hengshuang Zhao

Simulation offers a scalable and efficient alternative to real-world data collection for learning visuomotor robotic policies. However, the simulation-to-reality, or Sim2Real distribution shift -- introduced by employing simulation-trained…

Robotics · Computer Science 2025-09-09 Yash Yardi , Samuel Biruduganti , Lars Ankile

Closed-loop evaluation is increasingly critical for end-to-end autonomous driving. Current closed-loop benchmarks using the CARLA simulator rely on manually configured traffic scenarios, which can diverge from real-world conditions,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Haibao Yu , Wenxian Yang , Ruiyang Hao , Chuanye Wang , Jiaru Zhong , Ping Luo , Zaiqing Nie

Simulating soft robots in cluttered environments remains an open problem due to the challenge of capturing complex dynamics and interactions with the environment. Furthermore, fast simulation is desired for quickly exploring robot behaviors…

Robotics · Computer Science 2020-11-04 Rianna Jitosho , Nathaniel Agharese , Allison Okamura , Zac Manchester

The robotics field is evolving towards data-driven, end-to-end learning, inspired by multimodal large models. However, reliance on expensive real-world data limits progress. Simulators offer cost-effective alternatives, but the gap between…

Robotics · Computer Science 2025-12-23 Hongwei Fan , Hang Dai , Jiyao Zhang , Jinzhou Li , Qiyang Yan , Yujie Zhao , Mingju Gao , Jinghang Wu , Hao Tang , Hao Dong

Benchmarking vision-based driving policies is challenging. On one hand, open-loop evaluation with real data is easy, but these results do not reflect closed-loop performance. On the other, closed-loop evaluation is possible in simulation,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Daniel Dauner , Marcel Hallgarten , Tianyu Li , Xinshuo Weng , Zhiyu Huang , Zetong Yang , Hongyang Li , Igor Gilitschenski , Boris Ivanovic , Marco Pavone , Andreas Geiger , Kashyap Chitta

Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as robots remains a significant challenge. This report…

Robotics · Computer Science 2025-03-27 Gemini Robotics Team , Saminda Abeyruwan , Joshua Ainslie , Jean-Baptiste Alayrac , Montserrat Gonzalez Arenas , Travis Armstrong , Ashwin Balakrishna , Robert Baruch , Maria Bauza , Michiel Blokzijl , Steven Bohez , Konstantinos Bousmalis , Anthony Brohan , Thomas Buschmann , Arunkumar Byravan , Serkan Cabi , Ken Caluwaerts , Federico Casarini , Oscar Chang , Jose Enrique Chen , Xi Chen , Hao-Tien Lewis Chiang , Krzysztof Choromanski , David D'Ambrosio , Sudeep Dasari , Todor Davchev , Coline Devin , Norman Di Palo , Tianli Ding , Adil Dostmohamed , Danny Driess , Yilun Du , Debidatta Dwibedi , Michael Elabd , Claudio Fantacci , Cody Fong , Erik Frey , Chuyuan Fu , Marissa Giustina , Keerthana Gopalakrishnan , Laura Graesser , Leonard Hasenclever , Nicolas Heess , Brandon Hernaez , Alexander Herzog , R. Alex Hofer , Jan Humplik , Atil Iscen , Mithun George Jacob , Deepali Jain , Ryan Julian , Dmitry Kalashnikov , M. Emre Karagozler , Stefani Karp , Chase Kew , Jerad Kirkland , Sean Kirmani , Yuheng Kuang , Thomas Lampe , Antoine Laurens , Isabel Leal , Alex X. Lee , Tsang-Wei Edward Lee , Jacky Liang , Yixin Lin , Sharath Maddineni , Anirudha Majumdar , Assaf Hurwitz Michaely , Robert Moreno , Michael Neunert , Francesco Nori , Carolina Parada , Emilio Parisotto , Peter Pastor , Acorn Pooley , Kanishka Rao , Krista Reymann , Dorsa Sadigh , Stefano Saliceti , Pannag Sanketi , Pierre Sermanet , Dhruv Shah , Mohit Sharma , Kathryn Shea , Charles Shu , Vikas Sindhwani , Sumeet Singh , Radu Soricut , Jost Tobias Springenberg , Rachel Sterneck , Razvan Surdulescu , Jie Tan , Jonathan Tompson , Vincent Vanhoucke , Jake Varley , Grace Vesom , Giulia Vezzani , Oriol Vinyals , Ayzaan Wahid , Stefan Welker , Paul Wohlhart , Fei Xia , Ted Xiao , Annie Xie , Jinyu Xie , Peng Xu , Sichun Xu , Ying Xu , Zhuo Xu , Yuxiang Yang , Rui Yao , Sergey Yaroshenko , Wenhao Yu , Wentao Yuan , Jingwei Zhang , Tingnan Zhang , Allan Zhou , Yuxiang Zhou

Driving simulation plays a crucial role in developing reliable driving agents by providing controlled, evaluative environments. To enable meaningful assessments, a high-quality driving simulator must satisfy several key requirements:…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Junzhe Jiang , Nan Song , Jingyu Li , Xiatian Zhu , Li Zhang

The advancement of embodied intelligence is accelerating the integration of robots into daily life as human assistants. This evolution requires robots to not only interpret high-level instructions and plan tasks but also perceive and adapt…

Robotics · Computer Science 2025-08-19 Zhichen Lou , Kechun Xu , Zhongxiang Zhou , Rong Xiong

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent…