中文
相关论文

相关论文: ORV: 4D Occupancy-centric Robot Video Generation

200 篇论文

3D occupancy prediction is an important task for the robustness of vision-centric autonomous driving, which aims to predict whether each point is occupied in the surrounding 3D space. Existing methods usually require 3D occupancy labels to…

计算机视觉与模式识别 · 计算机科学 2023-11-30 Yuanhui Huang , Wenzhao Zheng , Borui Zhang , Jie Zhou , Jiwen Lu

Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes,…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Zeyi Liu , Shuang Li , Eric Cousineau , Siyuan Feng , Benjamin Burchfiel , Shuran Song

Videos depict the change of complex dynamical systems over time in the form of discrete image sequences. Generating controllable videos by learning the dynamical system is an important yet underexplored topic in the computer vision…

计算机视觉与模式识别 · 计算机科学 2023-04-05 Yucheng Xu , Li Nanbo , Arushi Goel , Zijian Guo , Zonghai Yao , Hamidreza Kasaei , Mohammadreze Kasaei , Zhibin Li

In recent years, autonomous driving has garnered escalating attention for its potential to relieve drivers' burdens and improve driving safety. Vision-based 3D occupancy prediction, which predicts the spatial occupancy status and semantics…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Yanan Zhang , Jinqing Zhang , Zengran Wang , Junhao Xu , Di Huang

Understanding the evolution of 3D scenes is important for effective autonomous driving. While conventional methods mode scene development with the motion of individual instances, world models emerge as a generative framework to describe the…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Lening Wang , Wenzhao Zheng , Yilong Ren , Han Jiang , Zhiyong Cui , Haiyang Yu , Jiwen Lu

Closed-loop simulation is essential for advancing end-to-end autonomous driving systems. Contemporary sensor simulation methods, such as NeRF and 3DGS, rely predominantly on conditions closely aligned with training data distributions, which…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Guosheng Zhao , Chaojun Ni , Xiaofeng Wang , Zheng Zhu , Xueyang Zhang , Yida Wang , Guan Huang , Xinze Chen , Boyuan Wang , Youyi Zhang , Wenjun Mei , Xingang Wang

Humanoid robot technology is advancing rapidly, with manufacturers introducing diverse heterogeneous visual perception modules tailored to specific scenarios. Among various perception paradigms, occupancy-based representation has become…

Event cameras offer microsecond latency, high dynamic range, and low power consumption, making them ideal for real-time robotic perception under challenging conditions such as motion blur, occlusion, and illumination changes. However,…

机器人学 · 计算机科学 2025-08-26 Krishna Vinod , Prithvi Jai Ramesh , Pavan Kumar B N , Bharatesh Chakravarthi

Occlusion is a longstanding difficulty that challenges the UAV-based object detection. Many works address this problem by adapting the detection model. However, few of them exploit that the UAV could fundamentally improve detection…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Xinhua Jiang , Tianpeng Liu , Li Liu , Zhen Liu , Yongxiang Liu

With the development of embodied artificial intelligence, robotic research has increasingly focused on complex tasks. Existing simulation platforms, however, are often limited to idealized environments, simple task scenarios and lack data…

机器人学 · 计算机科学 2025-04-29 Zijie Zheng , Zeshun Li , Yunpeng Wang , Qinghongbing Xie , Long Zeng

Video generative models are increasingly used as world models for robotics, where a model generates a future visual rollout conditioned on the current observation and task instruction, and an inverse dynamics model (IDM) converts the…

机器人学 · 计算机科学 2026-03-25 Ruixiang Wang , Qingming Liu , Yueci Deng , Guiliang Liu , Zhen Liu , Kui Jia

Single-stream architectures using Vision Transformer (ViT) backbones show great potential for real-time UAV tracking recently. However, frequent occlusions from obstacles like buildings and trees expose a major drawback: these models often…

计算机视觉与模式识别 · 计算机科学 2025-04-15 You Wu , Xucheng Wang , Xiangyang Yang , Mengyuan Liu , Dan Zeng , Hengzhou Ye , Shuiwang Li

The deployment of humanoid robots for dexterous manipulation in unstructured environments remains challenging due to perceptual limitations that constrain the effective workspace. In scenarios where physical constraints prevent the robot…

机器人学 · 计算机科学 2026-03-09 Pei Qu , Zheng Li , Yufei Jia , Ziyun Liu , Liang Zhu , Haoang Li , Jinni Zhou , Jun Ma

Robust 3D occupancy prediction is essential for autonomous driving, particularly under adverse weather conditions where traditional vision-only systems struggle. While the fusion of surround-view 4D radar and cameras offers a promising…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Long Yang , Lianqing Zheng , Wenjin Ai , Minghao Liu , Sen Li , Qunshu Lin , Shengyu Yan , Jie Bai , Zhixiong Ma , Tao Huang , Xichan Zhu

In this work, we present Patch-based Object-centric Video Transformer (POVT), a novel region-based video generation architecture that leverages object-centric information to efficiently model temporal dynamics in videos. We build upon prior…

计算机视觉与模式识别 · 计算机科学 2022-06-22 Wilson Yan , Ryo Okumura , Stephen James , Pieter Abbeel

3D occupancy-based perception pipeline has significantly advanced autonomous driving by capturing detailed scene descriptions and demonstrating strong generalizability across various object categories and shapes. Current methods…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Fangqiang Ding , Xiangyu Wen , Yunzhou Zhu , Yiming Li , Chris Xiaoxuan Lu

Visual Active Tracking (VAT) aims to control cameras to follow a target in 3D space, which is critical for applications like drone navigation and security surveillance. However, it faces two key bottlenecks in real-world deployment:…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Haowei Sun , Kai Zhou , Hao Gao , Shiteng Zhang , Jinwu Hu , Xutao Wen , Qixiang Ye , Mingkui Tan

Fast, collision-free motion through unknown environments remains a challenging problem for robotic systems. In these situations, the robot's ability to reason about its future motion is often severely limited by sensor field of view (FOV).…

机器学习 · 计算机科学 2018-03-07 Kapil Katyal , Katie Popek , Chris Paxton , Joseph Moore , Kevin Wolfe , Philippe Burlina , Gregory D. Hager

To effectively apply robots in working environments and assist humans, it is essential to develop and evaluate how visual grounding (VG) can affect machine performance on occluded objects. However, current VG works are limited in working…

计算与语言 · 计算机科学 2021-04-15 Ke-Jyun Wang , Yun-Hsuan Liu , Hung-Ting Su , Jen-Wei Wang , Yu-Siang Wang , Winston H. Hsu , Wen-Chin Chen

The field of autonomous driving is experiencing a surge of interest in world models, which aim to predict potential future scenarios based on historical observations. In this paper, we introduce DFIT-OccWorld, an efficient 3D occupancy…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Haiming Zhang , Ying Xue , Xu Yan , Jiacheng Zhang , Weichao Qiu , Dongfeng Bai , Bingbing Liu , Shuguang Cui , Zhen Li