中文
相关论文

相关论文: PanoWorld: Geometry-Consistent Panoramic Video Wor…

200 篇论文

Omnidirectional images are increasingly used in robotics and vision due to their wide field of view. However, extending 3D Gaussian Splatting (3DGS) to panoramic camera models remains challenging, as existing formulations are designed for…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Zhe Yang , Guoqiang Zhao , Sheng Wu , Kai Luo , Kailun Yang

AI video generation has lowered barriers to video creation, but current tools still struggle with inconsistency. Filmmakers often find that clips fail to match characters and backgrounds, making it difficult to build coherent sequences. A…

人机交互 · 计算机科学 2025-12-22 Hye-Young Jo , Mose Sakashita , Aditi Mishra , Ryo Suzuki , Koichiro Niinuma , Aakar Gupta

We present a novel algorithm for estimating the broad 3D geometric structure of outdoor video scenes. Leveraging spatio-temporal video segmentation, we decompose a dynamic scene captured by a video into geometric classes, based on…

计算机视觉与模式识别 · 计算机科学 2016-11-17 S. Hussain Raza , Matthias Grundmann , Irfan Essa

As Artificial Intelligence Generated Content (AIGC) advances, a variety of methods have been developed to generate text, images, videos, and 3D objects from single or multimodal inputs, contributing efforts to emulate human-like cognitive…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Yiying Yang , Fukun Yin , Jiayuan Fan , Xin Chen , Wanzhang Li , Gang Yu

In this work, we propose DiT360, a DiT-based framework that performs hybrid training on perspective and panoramic data for panoramic image generation. For the issues of maintaining geometric fidelity and photorealism in generation quality,…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Haoran Feng , Dizhe Zhang , Xiangtai Li , Bo Du , Lu Qi

3D scene generation is a core technology for gaming, film/VFX, and VR/AR. Growing demand for rapid iteration, high-fidelity detail, and accessible content creation has further increased interest in this area. Existing methods broadly follow…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Haozhi Zhu , Miaomiao Zhao , Dingyao Liu , Runze Tian , Yan Zhang , Jie Guo , Fenggen Yu

Towards building comprehensive real-world visual perception systems, we propose and study a new problem called panoptic scene graph generation (PVSG). PVSG relates to the existing video scene graph generation (VidSGG) problem, which focuses…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Jingkang Yang , Wenxuan Peng , Xiangtai Li , Zujin Guo , Liangyu Chen , Bo Li , Zheng Ma , Kaiyang Zhou , Wayne Zhang , Chen Change Loy , Ziwei Liu

Generating flexible-view 3D scenes, including 360{\deg} rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework consisting of two key components: (1) a strong…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Luxi Chen , Zihan Zhou , Min Zhao , Yikai Wang , Ge Zhang , Wenhao Huang , Hao Sun , Ji-Rong Wen , Chongxuan Li

3D Gaussian Splatting (3DGS) delivers high-fidelity real-time rendering but suffers from geometric and photometric degradations under sparse-view constraints. Current generative restoration approaches are often limited by insufficient…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Xinliang Wang , Yifeng Shi , Zhenyu Wu

Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated videos still contain visual artifacts arising from 3D…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Duolikun Danier , Ge Gao , Steven McDonagh , Changjian Li , Hakan Bilen , Oisin Mac Aodha

Wide-baseline panoramic images are frequently used in applications like VR and simulations to minimize capturing labor costs and storage needs. However, synthesizing novel views from these panoramic images in real time remains a significant…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Zheng Chen , Chenming Wu , Zhelun Shen , Chen Zhao , Weicai Ye , Haocheng Feng , Errui Ding , Song-Hai Zhang

World models that forecast scene evolution by generating future video frames devote the bulk of their capacity to photometric details, yet the resulting predictions often remain geometrically inconsistent. We present VGGT-World, a geometry…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Xiangyu Sun , Shijie Wang , Fengyi Zhang , Lin Liu , Caiyan Jia , Ziying Song , Zi Huang , Yadan Luo

This paper presents WorldPlay, a streaming video diffusion model that enables real-time, interactive world modeling with long-term geometric consistency, resolving the trade-off between speed and memory that limits current methods.…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Wenqiang Sun , Haiyu Zhang , Haoyuan Wang , Junta Wu , Zehan Wang , Zhenwei Wang , Yunhong Wang , Jun Zhang , Tengfei Wang , Chunchao Guo

The presence of spherical distortion in equirectangular projection (ERP) images presents a persistent challenge in dense regression tasks such as surface normal estimation. Although it may appear straightforward to repurpose architectures…

计算机视觉与模式识别 · 计算机科学 2026-01-26 Kun Huang , Fanglue Zhang , Neil Dodgson

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibilities for immersive…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Taewoong Kang , Kinam Kim , Dohyeon Kim , Minho Park , Junha Hyung , Jaegul Choo

In this paper, we tackle the problem of synthesizing a ground-view panorama image conditioned on a top-view aerial image, which is a challenging problem due to the large gap between the two image domains with different view-points. Instead…

计算机视觉与模式识别 · 计算机科学 2022-03-23 Songsong Wu , Hao Tang , Xiao-Yuan Jing , Haifeng Zhao , Jianjun Qian , Nicu Sebe , Yan Yan

The field of autonomous driving increasingly demands high-quality annotated video training data. In this paper, we propose Panacea+, a powerful and universally applicable framework for generating video data in driving scenes. Built upon the…

计算机视觉与模式识别 · 计算机科学 2024-08-15 Yuqing Wen , Yucheng Zhao , Yingfei Liu , Binyuan Huang , Fan Jia , Yanhui Wang , Chi Zhang , Tiancai Wang , Xiaoyan Sun , Xiangyu Zhang

Recent progress in image and video synthesis has inspired their use in advancing 3D scene generation. However, we observe that text-to-image and -video approaches struggle to maintain scene- and object-level consistency beyond a limited…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Manuel-Andreas Schneider , Angela Dai

Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First, they exhibit motion drift in complex environments with…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Guangyuan Li , Bo Li , Jinwei Chen , Xiaobin Hu , Lei Zhao , Peng-Tao Jiang

Recent successes in autoregressive (AR) generation models, such as the GPT series in natural language processing, have motivated efforts to replicate this success in visual tasks. Some works attempt to extend this approach to autonomous…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Xiaotao Hu , Wei Yin , Mingkai Jia , Junyuan Deng , Xiaoyang Guo , Qian Zhang , Xiaoxiao Long , Ping Tan