中文
相关论文

相关论文: Captain Safari: A World Engine with Pose-Aligned 3…

200 篇论文

Embodied navigation in open, dynamic environments demands accurate foresight of how the world will evolve and how actions will unfold over time. We propose AstraNav-World, an end-to-end world model that jointly reasons about future visual…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Jintao Chen , Junjun Hu , Haochen Bai , Minghua Luo , Xinda Xue , Botao Ren , Chengyu Bai , Shichao Xie , Ziyi Chen , Fei Liu , Zedong Chu , Xiaolong Wu , Mu Xu , Shanghang Zhang

Existing automatic approaches for 3D virtual character motion synthesis supporting scene interactions do not generalise well to new objects outside training distributions, even when trained on extensive motion capture datasets with diverse…

计算机视觉与模式识别 · 计算机科学 2024-02-16 Wanyue Zhang , Rishabh Dabral , Thomas Leimkühler , Vladislav Golyanik , Marc Habermann , Christian Theobalt

This study seeks to automate camera movement control for filming existing subjects into attractive videos, contrasting with the creation of non-existent content by directly generating the pixels. We select drone videos as our test case due…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Yunzhong Hou , Liang Zheng , Philip Torr

Search and rescue environments exhibit challenging 3D geometry (e.g., confined spaces, rubble, and breakdown), which necessitates agile and maneuverable aerial robotic systems. Because these systems are size, weight, and power (SWaP)…

机器人学 · 计算机科学 2024-11-08 Jonathan Lee , Abhishek Rathod , Kshitij Goel , John Stecklein , Wennie Tabib

Existing deep learning approaches on 3d human pose estimation for videos are either based on Recurrent or Convolutional Neural Networks (RNNs or CNNs). However, RNN-based frameworks can only tackle sequences with limited frames because…

计算机视觉与模式识别 · 计算机科学 2019-08-23 Jiahao Lin , Gim Hee Lee

World models and video generation are pivotal technologies in the domain of autonomous driving, each playing a critical role in enhancing the robustness and reliability of autonomous systems. World models, which simulate the dynamics of…

人工智能 · 计算机科学 2024-11-06 Ao Fu , Yi Zhou , Tao Zhou , Yi Yang , Bojun Gao , Qun Li , Guobin Wu , Ling Shao

Pose tracking of uncooperative spacecraft is an essential technology for space exploration and on-orbit servicing, which remains an open problem. Event cameras possess numerous advantages, such as high dynamic range, high temporal…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Zibin Liu , Banglei Guan , Yang Shang , Yifei Bian , Pengju Sun , Qifeng Yu

Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models forecast the external environment, in-cabin intelligence remains strictly…

机器人学 · 计算机科学 2026-05-07 Haozhuang Chi , Daosheng Qiu , Hao Su , Haochen Liu , Zirui Li , Haoruo Zhang , Chen Lv

Storytelling in real-world videos often unfolds through multiple shots -- discontinuous yet semantically connected clips that together convey a coherent narrative. However, existing multi-shot video generation (MSV) methods struggle to…

End-to-end autonomous driving directly generates planning trajectories from raw sensor data, yet it typically relies on costly perception supervision to extract scene information. A critical research challenge arises: constructing an…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Yupeng Zheng , Pengxuan Yang , Zebin Xing , Qichao Zhang , Yuhang Zheng , Yinfeng Gao , Pengfei Li , Teng Zhang , Zhongpu Xia , Peng Jia , Dongbin Zhao

Closed-loop simulation is crucial for end-to-end autonomous driving. Existing sensor simulation methods (e.g., NeRF and 3DGS) reconstruct driving scenes based on conditions that closely mirror training data distributions. However, these…

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Yue Gao , Yuan Zhou , Jinglu Wang , Xiao Li , Xiang Ming , Yan Lu

Recent video-based world models have made pixel-space environments interactive at the camera level: users can navigate viewpoints while the model generates coherent visual continuations. Yet their action spaces remain incomplete: users can…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Bohai Gu , Taiyi Wu , Yueyang Yuan , Jian Liu , Xiaocheng Lu , Dazhao Du , Jie Zhang , Jinxiang Lai , Shuai Yang , Xiaotong Zhao , Alan Zhao , Song Guo

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Tianze Xia , Yongkang Li , Lijun Zhou , Jingfeng Yao , Kaixin Xiong , Haiyang Sun , Bing Wang , Kun Ma , Guang Chen , Hangjun Ye , Wenyu Liu , Xinggang Wang

Autoregressive video generation paradigms offer theoretical promise for long video synthesis, yet their practical deployment is hindered by the computational burden of sequential iterative denoising. While cache reuse strategies can…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Jing Xu , Yuexiao Ma , Xuzhe Zheng , Xing Wang , Shiwei Liu , Chenqian Yan , Xiawu Zheng , Rongrong Ji , Fei Chao , Songwei Liu

A world model is an AI system that simulates how an environment evolves under actions, enabling planning through imagined futures rather than reactive perception. Current world models, however, suffer from visual conflation: the mistaken…

人工智能 · 计算机科学 2026-01-23 Zhikang Chen , Tingting Zhu

Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Taiye Chen , Xun Hu , Zihan Ding , Chi Jin

Recent progress in driving video generation has shown significant potential for enhancing self-driving systems by providing scalable and controllable training data. Although pretrained state-of-the-art generation models, guided by 2D layout…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Yishen Ji , Ziyue Zhu , Zhenxin Zhu , Kaixin Xiong , Ming Lu , Zhiqi Li , Lijun Zhou , Haiyang Sun , Bing Wang , Tong Lu

Recently, breakthroughs in video modeling have allowed for controllable camera trajectories in generated videos. However, these methods cannot be directly applied to user-provided videos that are not generated by a video model. In this…

计算机视觉与模式识别 · 计算机科学 2024-11-08 David Junhao Zhang , Roni Paiss , Shiran Zada , Nikhil Karnad , David E. Jacobs , Yael Pritch , Inbar Mosseri , Mike Zheng Shou , Neal Wadhwa , Nataniel Ruiz

Six degree of freedom (6DoF) pose estimation for novel objects is a critical task in computer vision, yet it faces significant challenges in high-speed and low-light scenarios where standard RGB cameras suffer from motion blur. While event…

计算机视觉与模式识别 · 计算机科学 2026-01-05 Huiming Yang , Linglin Liao , Fei Ding , Sibo Wang , Zijian Zeng