中文
相关论文

相关论文: Cross-View World Models

200 篇论文

We propose improving the cross-target and cross-scene generalization of visual navigation through learning an agent that is guided by conceiving the next observations it expects to see. This is achieved by learning a variational Bayesian…

机器人学 · 计算机科学 2022-01-11 Qiaoyun Wu , Dinesh Manocha , Jun Wang , Kai Xu

Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scenes evolve under egocentric motion. Recent efforts address…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Wanyue Zhang , Wenxiang Wu , Wang Xu , Jiaxin Luo , Helu Zhi , Yibin Huang , Shuo Ren , Zitao Liu , Jiajun Zhang

Cross-view geo-localization is the problem of estimating the position and orientation (latitude, longitude and azimuth angle) of a camera at ground level given a large-scale database of geo-tagged aerial (e.g., satellite) images. Existing…

计算机视觉与模式识别 · 计算机科学 2020-05-11 Yujiao Shi , Xin Yu , Dylan Campbell , Hongdong Li

Embodied systems experience the world as 'a symphony of flows': a combination of many continuous streams of sensory input coupled to self-motion, interwoven with the dynamics of external objects. These streams obey smooth,…

机器学习 · 计算机科学 2026-01-06 Hansen Jin Lillemark , Benhao Huang , Fangneng Zhan , Yilun Du , Thomas Anderson Keller

High-quality 3D world models are pivotal for embodied intelligence and Artificial General Intelligence (AGI), underpinning applications such as AR/VR content creation and robotic navigation. Despite the established strong imaginative…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Yixiang Dai , Fan Jiang , Chiyu Wang , Mu Xu , Yonggang Qi

Convolutions on monocular dash cam videos capture spatial invariances in the image plane but do not explicitly reason about distances and depth. We propose a simple transformation of observations into a bird's eye view, also known as plan…

计算机视觉与模式识别 · 计算机科学 2019-05-17 Dequan Wang , Coline Devin , Qi-Zhi Cai , Philipp Krähenbühl , Trevor Darrell

We explore building generative neural network models of popular reinforcement learning environments. Our world model can be trained quickly in an unsupervised manner to learn a compressed spatial and temporal representation of the…

机器学习 · 计算机科学 2018-05-10 David Ha , Jürgen Schmidhuber

World models learn the consequences of actions in vision-based interactive systems. However, in practical scenarios such as autonomous driving, there commonly exists noncontrollable dynamics independent of the action signals, making it…

机器学习 · 计算机科学 2022-12-20 Minting Pan , Xiangming Zhu , Yunbo Wang , Xiaokang Yang

World models aim to capture the dynamics of the environment, enabling agents to predict and plan for future states. In most scenarios of interest, the dynamics are highly centered on interactions among objects within the environment. This…

机器学习 · 计算机科学 2025-07-08 Zizhao Wang , Kaixin Wang , Li Zhao , Peter Stone , Jiang Bian

Learning predictive models from interaction with the world allows an agent, such as a robot, to learn about how the world works, and then use this learned model to plan coordinated sequences of actions to bring about desired outcomes.…

机器学习 · 计算机科学 2020-01-01 Karl Schmeckpeper , Annie Xie , Oleh Rybkin , Stephen Tian , Kostas Daniilidis , Sergey Levine , Chelsea Finn

In this paper, we investigate the problem of embodied multi-agent cooperation, where decentralized agents must cooperate given only egocentric views of the world. To effectively plan in this setting, in contrast to learning world dynamics…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Hongxin Zhang , Zeyuan Wang , Qiushi Lyu , Zheyuan Zhang , Sunli Chen , Tianmin Shu , Behzad Dariush , Kwonjoon Lee , Yilun Du , Chuang Gan

The challenge of navigation in environments with dynamic objects continues to be a central issue in the study of autonomous agents. While predictive methods hold promise, their reliance on precise state information makes them less practical…

机器人学 · 计算机科学 2024-10-28 Hsuan-Kung Yang , Tsung-Chih Chiang , Ting-Ru Liu , Chun-Wei Huang , Jou-Min Liu , Chun-Yi Lee

Ophthalmic decision-making depends on subtle lesion-scale cues interpreted across multimodal imaging and over time, yet most medical foundation models remain static and degrade under modality and acquisition shifts. Here we introduce…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Ziyu Gao , Xinyuan Wu , Xiaolan Chen , Zhuoran Liu , Ruoyu Chen , Bowen Liu , Bingjie Yan , Zhenhan Wang , Kai Jin , Jiancheng Yang , Yih Chung Tham , Mingguang He , Danli Shi

Humans are excellent at understanding language and vision to accomplish a wide range of tasks. In contrast, creating general instruction-following embodied agents remains a difficult challenge. Prior work that uses pure language-only models…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Hao Liu , Lisa Lee , Kimin Lee , Pieter Abbeel

Trajectory prediction and planning are essential for autonomous vehicles to navigate safely and efficiently in dynamic environments. Traditional approaches often treat them separately, limiting the ability for interactive planning. While…

机器人学 · 计算机科学 2025-07-22 Anjian Li , Sangjae Bae , David Isele , Ryne Beeson , Faizan M. Tariq

A long-standing challenge in Reinforcement Learning is enabling agents to learn a model of their environment which can be transferred to solve other problems in a world with the same underlying rules. One reason this is difficult is the…

机器学习 · 计算机科学 2019-05-16 Kai Olav Ellefsen , Jim Torresen

World Models have vastly permeated the field of Reinforcement Learning. Their ability to model the transition dynamics of an environment have greatly improved sample efficiency in online RL. Among them, the most notorious example is…

机器学习 · 计算机科学 2025-10-21 Federico Malato , Ville Hautamäki

General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness remains challenging. Unlike policy learning, which primarily focuses on optimal actions, a world…

机器学习 · 计算机科学 2026-04-03 Yuejiang Liu , Fan Feng , Lingjing Kong , Weifeng Lu , Jinzhou Tang , Kun Zhang , Kevin Murphy , Chelsea Finn , Yilun Du

We study language-conditioned visual navigation (LCVN), in which an embodied agent is asked to follow a natural language instruction based only on an initial egocentric observation. Without access to goal images, the agent must rely on…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yifei Dong , Fengyi Wu , Yilong Dai , Lingdong Kong , Guangyu Chen , Xu Zhu , Qiyu Hu , Tianyu Wang , Johnalbert Garnica , Feng Liu , Siyu Huang , Qi Dai , Zhi-Qi Cheng

Recent advances in video diffusion transformers have enabled interactive gaming world models that allow users to explore generated environments over extended horizons. However, existing approaches struggle with precise action control and…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Jisu Nam , Yicong Hong , Chun-Hao Paul Huang , Feng Liu , JoungBin Lee , Jiyoung Kim , Siyoon Jin , Yunsung Lee , Jaeyoon Jung , Suhwan Choi , Seungryong Kim , Yang Zhou