中文
相关论文

相关论文: PanoWorld: A Generative Spatial World Model for Co…

200 篇论文

Achieving an immersive experience enabling users to explore virtual environments with six degrees of freedom (6DoF) is essential for various applications such as virtual reality (VR). Wide-baseline panoramas are commonly used in these…

计算机视觉与模式识别 · 计算机科学 2023-12-07 Zheng Chen , Yan-Pei Cao , Yuan-Chen Guo , Chen Wang , Ying Shan , Song-Hai Zhang

Generating explorable 3D scenes from a single image is a highly challenging problem in 3D vision. Existing methods struggle to support free exploration, often producing severe geometric distortions and noisy artifacts when the viewpoint…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Pengfei Wang , Liyi Chen , Zhiyuan Ma , Yanjun Guo , Guowen Zhang , Lei Zhang

The blooming of virtual reality and augmented reality (VR/AR) technologies has driven an increasing demand for the creation of high-quality, immersive, and dynamic environments. However, existing generative techniques either focus solely on…

计算机视觉与模式识别 · 计算机科学 2024-10-04 Renjie Li , Panwang Pan , Bangbang Yang , Dejia Xu , Shijie Zhou , Xuanyang Zhang , Zeming Li , Achuta Kadambi , Zhangyang Wang , Zhengzhong Tu , Zhiwen Fan

3D Visual Grounding (3DVG) is a critical bridge from vision-language perception to robotics, requiring both language understanding and 3D scene reasoning. Traditional supervised models leverage explicit 3D geometry but exhibit limited…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Seongmin Jung , Seongho Choi , Gunwoo Jeon , Minsu Cho , Jongwoo Lim

Achieving realistic hair strand synthesis is essential for creating lifelike digital humans, but producing high-fidelity hair strand geometry remains a significant challenge. Existing methods require a complex setup for data acquisition,…

图形学 · 计算机科学 2025-08-27 Shashikant Verma , Shanmuganathan Raman

Recent advances in image generation have led to remarkable improvements in synthesizing perspective images. However, these models still struggle with panoramic image generation due to unique challenges, including varying levels of geometric…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Hakan Çapuk , Andrew Bond , Muhammed Burak Kızıl , Emir Göçen , Erkut Erdem , Aykut Erdem

We propose FlashWorld, a generative model that produces 3D scenes from a single image or text prompt in seconds, 10~100$\times$ faster than previous works while possessing superior rendering quality. Our approach shifts from the…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Xinyang Li , Tengfei Wang , Zixiao Gu , Shengchuan Zhang , Chunchao Guo , Liujuan Cao

The volume and diversity of training data are critical for modern deep learningbased methods. Compared to the massive amount of labeled perspective images, 360 panoramic images fall short in both volume and diversity. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2023-09-28 Yu-Cheng Hsieh , Cheng Sun , Suraj Dengale , Min Sun

3D panorama synthesis is a promising yet challenging task that demands high-quality and diverse visual appearance and geometry of the generated omnidirectional content. Existing methods leverage rich image priors from pre-trained 2D…

图形学 · 计算机科学 2025-06-23 Yukun Huang , Yanning Zhou , Jianan Wang , Kaiyi Huang , Xihui Liu

We present GuidedSceneGen, a text-to-3D generation framework that produces metrically accurate, globally consistent, and semantically interpretable indoor scenes. Unlike prior text-driven methods that often suffer from geometric drift or…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Stefan Ainetter , Thomas Deixelberger , Edoardo A. Dominici , Philipp Drescher , Konstantinos Vardis , Markus Steinberger

We tackle the challenge of generating the infinitely extendable 3D world -- large, continuous environments with coherent geometry and realistic appearance. Existing methods face key challenges: 2D-lifting approaches suffer from geometric…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Sikuang Li , Chen Yang , Jiemin Fang , Taoran Yi , Jia Lu , Jiazhong Cen , Lingxi Xie , Wei Shen , Qi Tian

We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prompts, single-view images, multi-view images, and videos, and…

Existing diffusion-based 3D scene generation methods primarily operate in 2D image/video latent spaces, which makes maintaining cross-view appearance and geometric consistency inherently challenging. To bridge this gap, we present OneWorld,…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Sensen Gao , Zhaoqing Wang , Qihang Cao , Dongdong Yu , Changhu Wang , Tongliang Liu , Mingming Gong , Jiawang Bian

Panoramic image stitching provides a unified, wide-angle view of a scene that extends beyond the camera's field of view. Stitching frames of a panning video into a panoramic photograph is a well-understood problem for stationary scenes, but…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Jingwei Ma , Erika Lu , Roni Paiss , Shiran Zada , Aleksander Holynski , Tali Dekel , Brian Curless , Michael Rubinstein , Forrester Cole

We present WonderWorld, a novel framework for interactive 3D scene generation that enables users to interactively specify scene contents and layout and see the created scenes in low latency. The major challenge lies in achieving fast…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Hong-Xing Yu , Haoyi Duan , Charles Herrmann , William T. Freeman , Jiajun Wu

In modern interior design, the generation of personalized spaces frequently necessitates a delicate balance between rigid architectural structural constraints and specific stylistic preferences. However, existing multi-condition generative…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Lulu Chen , Yijiang Hu , Yuanqing Liu , Yulong Li , Yue Yang

In this paper, we tackle the problem of synthesizing a ground-view panorama image conditioned on a top-view aerial image, which is a challenging problem due to the large gap between the two image domains with different view-points. Instead…

计算机视觉与模式识别 · 计算机科学 2022-03-23 Songsong Wu , Hao Tang , Xiao-Yuan Jing , Haifeng Zhao , Jianjun Qian , Nicu Sebe , Yan Yan

Taking a picture has been traditionally a one-persons task. In this paper we present a novel system that allows multiple mobile devices to work collaboratively in a synchronized fashion to capture a panorama of a highly dynamic scene,…

人机交互 · 计算机科学 2015-07-13 Yan Wang , Sunghyun Cho , Jue Wang , Shih-Fu Chang

We study the problem of synthesizing immersive 3D indoor scenes from one or more images. Our aim is to generate high-resolution images and videos from novel viewpoints, including viewpoints that extrapolate far beyond the input images while…

计算机视觉与模式识别 · 计算机科学 2022-12-02 Jing Yu Koh , Harsh Agrawal , Dhruv Batra , Richard Tucker , Austin Waters , Honglak Lee , Yinfei Yang , Jason Baldridge , Peter Anderson

High-quality 3D world models are pivotal for embodied intelligence and Artificial General Intelligence (AGI), underpinning applications such as AR/VR content creation and robotic navigation. Despite the established strong imaginative…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Yixiang Dai , Fan Jiang , Chiyu Wang , Mu Xu , Yonggang Qi