中文
相关论文

相关论文: QuaDreamer: Controllable Panoramic Video Generatio…

200 篇论文

As virtual reality gains popularity, the demand for controllable creation of immersive and dynamic omnidirectional videos (ODVs) is increasing. While previous text-to-ODV generation methods achieve impressive results, they struggle with…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Weiqi Li , Shijie Zhao , Chong Mou , Xuhan Sheng , Zhenyu Zhang , Qian Wang , Junlin Li , Li Zhang , Jian Zhang

Traditional approaches to quadruped control frequently employ simplified, hand-derived models. This significantly reduces the capability of the robot since its effective kinematic range is curtailed. In addition, kinodynamic constraints are…

360{\deg} videos have emerged as a promising medium to represent our dynamic visual world. Compared to the "tunnel vision" of standard cameras, their borderless field of view offers a more complete perspective of our surroundings. While…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Rundong Luo , Matthew Wallingford , Ali Farhadi , Noah Snavely , Wei-Chiu Ma

Surround-view perception is increasingly important for robotic navigation and loco-manipulation, especially in human-in-the-loop settings such as teleoperation, data collection, and emergency takeover. However, current robotic visual…

Multimodal locomotion capability is an emerging topic in robotics field, and various novel mobile robots have been developed to enable the maneuvering in both terrestrial and aerial domains. Among these hybrid robots, several…

机器人学 · 计算机科学 2023-01-11 Moju Zhao , Tomoki Anzai , Takuzumi Nishio

Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes,…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Zeyi Liu , Shuang Li , Eric Cousineau , Siyuan Feng , Benjamin Burchfiel , Shuran Song

Quadruped robots are currently used in industrial robotics as mechanical aid to automate several routine tasks. However, presently, the usage of such a robot in a domestic setting is still very much a part of the research. This paper…

Multimodal audiovisual perception can enable new avenues for robotic manipulation, from better material classification to the imitation of demonstrations for which only audio signals are available (e.g., playing a tune by ear). However, to…

机器人学 · 计算机科学 2026-03-09 Luca Macesanu , Boueny Folefack , Samik Singh , Ruchira Ray , Ben Abbatematteo , Roberto Martín-Martín

A guiding robot aims to effectively bring people to and from specific places within environments that are possibly unknown to them. During this operation the robot should be able to detect and track the accompanied person, trying never to…

计算机视觉与模式识别 · 计算机科学 2023-07-31 Alberto Bacchin , Filippo Berno , Emanuele Menegatti , Alberto Pretto

The blooming of virtual reality and augmented reality (VR/AR) technologies has driven an increasing demand for the creation of high-quality, immersive, and dynamic environments. However, existing generative techniques either focus solely on…

计算机视觉与模式识别 · 计算机科学 2024-10-04 Renjie Li , Panwang Pan , Bangbang Yang , Dejia Xu , Shijie Zhou , Xuanyang Zhang , Zeming Li , Achuta Kadambi , Zhangyang Wang , Zhengzhong Tu , Zhiwen Fan

The field of autonomous driving increasingly demands high-quality annotated training data. In this paper, we propose Panacea, an innovative approach to generate panoramic and controllable videos in driving scenarios, capable of yielding an…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Yuqing Wen , Yucheng Zhao , Yingfei Liu , Fan Jia , Yanhui Wang , Chong Luo , Chi Zhang , Tiancai Wang , Xiaoyan Sun , Xiangyu Zhang

Previous Vision-Language-Action models face critical limitations in navigation: scarce, diverse data from labor-intensive collection and static representations that fail to capture temporal dynamics and physical laws. We propose NavDreamer,…

机器人学 · 计算机科学 2026-02-11 Xijie Huang , Weiqi Gai , Tianyue Wu , Congyu Wang , Zhiyang Liu , Xin Zhou , Yuze Wu , Fei Gao

Reliable onboard perception is critical for quadruped robots navigating dynamic environments, where obstacles can emerge from any direction under strict reaction-time constraints. Single-sensor systems face inherent limitations: LiDAR…

机器人学 · 计算机科学 2025-09-30 Zihao Xu , Kuankuan Sima , Junhao Deng , Zixuan Zhuang , Chunzheng Wang , Ce Hao , Jin Song Dong

Panoramic perception holds significant potential for autonomous driving, enabling vehicles to acquire a comprehensive 360{\deg} surround view in a single shot. However, autonomous driving is a data-driven task. Complete panoramic data…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Fei Teng , Kai Luo , Sheng Wu , Siyu Li , Pujun Guo , Jiale Wei , Jiaming Zhang , Kunyu Peng , Kailun Yang

3D scene generation is in high demand across various domains, including virtual reality, gaming, and the film industry. Owing to the powerful generative capabilities of text-to-image diffusion models that provide reliable priors, the…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Haiyang Zhou , Xinhua Cheng , Wangbo Yu , Yonghong Tian , Li Yuan

Quadrupedal robots are conquering various indoor and outdoor applications due to their ability to navigate challenging uneven terrains. Exteroceptive information greatly enhances this capability since perceiving their surroundings allows…

机器人学 · 计算机科学 2023-04-07 Benedek Forrai , Takahiro Miki , Daniel Gehrig , Marco Hutter , Davide Scaramuzza

We introduce VividDream, a method for generating explorable 4D scenes with ambient dynamics from a single input image or text prompt. VividDream first expands an input image into a static 3D point cloud through iterative inpainting and…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Yao-Chih Lee , Yi-Ting Chen , Andrew Wang , Ting-Hsuan Liao , Brandon Y. Feng , Jia-Bin Huang

Generative models have been widely applied to world modeling for environment simulation and future state prediction. With advancements in autonomous driving, there is a growing demand not only for high-fidelity video generation under…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Tianrui Zhang , Yichen Liu , Zilin Guo , Yuxin Guo , Jingcheng Ni , Chenjing Ding , Dan Xu , Lewei Lu , Zehuan Wu

In this work, we present SceneDreamer, an unconditional generative model for unbounded 3D scenes, which synthesizes large-scale 3D landscapes from random noise. Our framework is learned from in-the-wild 2D image collections only, without…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Zhaoxi Chen , Guangcong Wang , Ziwei Liu

This project aimed to develop an automated cinematography platform using an unmanned aerial vehicle. Quadcopters are a great platform for shooting aerial scenes but are difficult to maneuver smoothly and can require expertise to pilot. We…

机器人学 · 计算机科学 2024-09-04 Animesh Nema , Christopher Grontkowski , Derek Calzada , Sanjuksha Nirgude