English
Related papers

Related papers: RoboDreamer: Learning Compositional World Models f…

200 papers

Current text-to-3D methods excel at generating single objects but falter on compositional prompts. We argue this failure is fundamental to their optimization schedules, as simultaneous or iterative heuristics predictably collapse under a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Utkarsh Nath , Rajeev Goel , Rahul Khurana , Kyle Min , Mark Ollila , Pavan Turaga , Varun Jampani , Tejaswi Gowda

In this work we propose a novel end-to-end imitation learning approach which combines natural language, vision, and motion information to produce an abstract representation of a task, which in turn is used to synthesize specific motion…

Robotics · Computer Science 2019-11-27 Simon Stepputtis , Joseph Campbell , Mariano Phielipp , Chitta Baral , Heni Ben Amor

Humans are excellent at understanding language and vision to accomplish a wide range of tasks. In contrast, creating general instruction-following embodied agents remains a difficult challenge. Prior work that uses pure language-only models…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Hao Liu , Lisa Lee , Kimin Lee , Pieter Abbeel

Industrial robots are designed as general-purpose hardware with limited ability to adapt to changing task requirements or environments. Modular robots, on the other hand, offer flexibility and can be easily customized to suit diverse needs.…

Robotics · Computer Science 2024-03-05 Jonathan Külz , Matthias Althoff

Building generalist robots capable of performing functional grasping in everyday, open-world environments remains a significant challenge due to the vast diversity of objects and tasks. Existing methods are either constrained to narrow…

Robotics · Computer Science 2026-04-10 Chao Tang , Jiacheng Xu , Haofei Lu , Bolin Zou , Wenlong Dong , Hong Zhang , Danica Kragic

Cognitive planning is the structural decomposition of complex tasks into a sequence of future behaviors. In the computational setting, performing cognitive planning entails grounding plans and concepts in one or more modalities in order to…

Artificial Intelligence · Computer Science 2022-10-11 Maria Attarian , Advaya Gupta , Ziyi Zhou , Wei Yu , Igor Gilitschenski , Animesh Garg

We present Emu Video, a text-to-video generation model that factorizes the generation into two steps: first generating an image conditioned on the text, and then generating a video conditioned on the text and the generated image. We…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Rohit Girdhar , Mannat Singh , Andrew Brown , Quentin Duval , Samaneh Azadi , Sai Saketh Rambhatla , Akbar Shah , Xi Yin , Devi Parikh , Ishan Misra

When large vision-language models are applied to the field of robotics, they encounter problems that are simple for humans yet error-prone for models. Such issues include confusion between third-person and first-person perspectives and a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Baiyu Pan , Daqin Luo , Junpeng Yang , Jiyuan Wang , Yixuan Zhang , Hailin Shi , Jichao Jiao

For a general-purpose robot to operate in reality, executing a broad range of instructions across various environments is imperative. Central to the reinforcement learning and planning for such robotic agents is a generalizable reward…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Yanting Yang , Minghao Chen , Qibo Qiu , Jiahao Wu , Wenxiao Wang , Binbin Lin , Ziyu Guan , Xiaofei He

Recent advances in video diffusion models shows promise for generating robotic decision-making data, with trajectory conditions further enabling fine-grained control. However, existing methods primarily focus on individual object motion and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Xiao Fu , Xintao Wang , Xian Liu , Jianhong Bai , Runsen Xu , Pengfei Wan , Di Zhang , Dahua Lin

World models have emerged as a powerful paradigm for building interactive simulation environments, with recent video-based approaches demonstrating impressive progress in generating visually plausible dynamics. However, because these models…

Artificial Intelligence · Computer Science 2026-05-15 Hongyu Wang , Jingquan Wang , Bocheng Zou , Radu Serban , Dan Negrut

3D open-world environments with adversarial opponents remain a core challenge for reinforcement learning due to their vast state spaces. Effective reasoning representations are essential in such settings. While existing self-supervised…

Artificial Intelligence · Computer Science 2026-05-26 Yuanfei Xu , Lin Liu , Wengang Zhou , Mingxiao Feng , Houqiang Li

Video Diffusion Models (VDMs) have emerged as powerful generative tools, capable of synthesizing high-quality spatiotemporal content. Yet, their potential goes far beyond mere video generation. We argue that the training dynamics of VDMs,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Pablo Acuaviva , Aram Davtyan , Mariam Hassan , Sebastian Stapf , Ahmad Rahimi , Alexandre Alahi , Paolo Favaro

With the advancement of generative models, the synthesis of different sensory elements such as music, visuals, and speech has achieved significant realism. However, the approach to generate multi-sensory outputs has not been fully explored,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Minheng Ni , Chenfei Wu , Huaying Yuan , Zhengyuan Yang , Ming Gong , Lijuan Wang , Zicheng Liu , Wangmeng Zuo , Nan Duan

Long-horizon embodied planning is challenging because the world does not only change through an agent's actions: exogenous processes (e.g., water heating, dominoes cascading) unfold concurrently with the agent's actions. We propose a…

Vision-language models such as CLIP have shown impressive capabilities in encoding texts and images into aligned embeddings, enabling the retrieval of multimodal data in a shared embedding space. However, these embedding-based models still…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Timothy Ossowski , Ming Jiang , Junjie Hu

World models simulate future states of the world in response to different actions. They facilitate interactive content creation and provides a foundation for grounded, long-horizon reasoning. Current foundation models do not fully meet the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Jiannan Xiang , Guangyi Liu , Yi Gu , Qiyue Gao , Yuting Ning , Yuheng Zha , Zeyu Feng , Tianhua Tao , Shibo Hao , Yemin Shi , Zhengzhong Liu , Eric P. Xing , Zhiting Hu

Diffusion models are capable of generating photo-realistic images that combine elements which likely do not appear together in the training set, demonstrating the ability to \textit{compositionally generalize}. Nonetheless, the precise…

Artificial Intelligence · Computer Science 2024-10-14 Qiyao Liang , Ziming Liu , Mitchell Ostrow , Ila Fiete

This work highlights that video world modeling, alongside vision-language pre-training, establishes a fresh and independent foundation for robot learning. Intuitively, video world models provide the ability to imagine the near future by…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Lin Li , Qihang Zhang , Yiming Luo , Shuai Yang , Ruilin Wang , Fei Han , Mingrui Yu , Zelin Gao , Nan Xue , Xing Zhu , Yujun Shen , Yinghao Xu

Language is compositional; an instruction can express multiple relation constraints to hold among objects in a scene that a robot is tasked to rearrange. Our focus in this work is an instructable scene-rearranging framework that generalizes…

‹ Prev 1 4 5 6 7 8 10 Next ›