中文
相关论文

相关论文: PanoWorld: Geometry-Consistent Panoramic Video Wor…

200 篇论文

In this paper, we explore the overlooked challenge of stability and temporal consistency in interactive video generation, which synthesizes dynamic and controllable video worlds through interactive behaviors such as camera movements and…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Ying Yang , Zhengyao Lv , Tianlin Pan , Haofan Wang , Binxin Yang , Hubery Yin , Chen Li , Ziwei Liu , Chenyang Si

World models have become a central paradigm for learning predictive simulators that support generation, planning, and decision-making. Yet, despite rapid progress in industry-scale interactive video generation, the broader research…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Siqiao Huang , Partha Kaushik , Michael Chen , Hengkai Pan , Kaiwen Geng , Omar Chehab , Fernando Moreno-Pino , Max Simchowitz

In this paper, we propose PanoViT, a panorama vision transformer to estimate the room layout from a single panoramic image. Compared to CNN models, our PanoViT is more proficient in learning global information from the panoramic image for…

计算机视觉与模式识别 · 计算机科学 2022-12-26 Weichao Shen , Yuan Dong , Zonghao Chen , Zhengyi Zhao , Yang Gao , Zhu Liu

Generating explorable 3D scenes from a single image is a highly challenging problem in 3D vision. Existing methods struggle to support free exploration, often producing severe geometric distortions and noisy artifacts when the viewpoint…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Pengfei Wang , Liyi Chen , Zhiyuan Ma , Yanjun Guo , Guowen Zhang , Lei Zhang

Scalable sensor simulation is an important yet challenging open problem for safety-critical domains such as self-driving. Current works in image simulation either fail to be photorealistic or do not model the 3D environment and the dynamic…

计算机视觉与模式识别 · 计算机科学 2021-05-18 Yun Chen , Frieda Rong , Shivam Duggal , Shenlong Wang , Xinchen Yan , Sivabalan Manivasagam , Shangjie Xue , Ersin Yumer , Raquel Urtasun

World models simulate future states of the world in response to different actions. They facilitate interactive content creation and provides a foundation for grounded, long-horizon reasoning. Current foundation models do not fully meet the…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Jiannan Xiang , Guangyi Liu , Yi Gu , Qiyue Gao , Yuting Ning , Yuheng Zha , Zeyu Feng , Tianhua Tao , Shibo Hao , Yemin Shi , Zhengzhong Liu , Eric P. Xing , Zhiting Hu

Geometric consistency, i.e. the preservation of neighbourhoods, is a natural and strong prior in 3D shape matching. Geometrically consistent matchings are crucial for many downstream applications, such as texture transfer or statistical…

图形学 · 计算机科学 2025-07-30 Paul Roetzer , Florian Bernard

Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geometric consistency. However, existing methods typically rely on fragmented geometric…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Hong Jiang , Wensong Song , Zongxing Yang , Ruijie Quan , Yi Yang

Wide-baseline panorama reconstruction has emerged as a highly effective and pivotal approach for not only achieving geometric reconstruction of the surrounding 3D environment, but also generating highly realistic and immersive novel views.…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Jiahui Ren , Mochu Xiang , Jiajun Zhu , Yuchao Dai

This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-independent training-free approaches, VGGT-360 reformulates the task as panoramic reprojection…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Jiayi Yuan , Haobo Jiang , De Wen Soh , Na Zhao

Video Diffusion Models (VDMs) offer a promising approach for simulating dynamic scenes and environments, with broad applications in robotics and media generation. However, existing models often generate temporally incoherent content that…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zhexiao Xiong , Yizhi Song , Liu He , Wei Xiong , Yu Yuan , Feng Qiao , Nathan Jacobs

We introduce CinemaWorld, a generative augmented reality system that augments the viewer's physical surroundings with automatically generated mixed reality 3D content extracted from and synchronized with 2D movie scenes. Our system…

人机交互 · 计算机科学 2026-03-10 Keiichi Ihara , DaeHo Lee , Manato Abe , Hye-Young Jo , Ryo Suzuki

Autonomous collision-free navigation in cluttered environments requires safe decision-making under partial observability with both static structure and dynamic obstacles. We present \textbf{PanoDP}, a communication-free learning framework…

机器人学 · 计算机科学 2026-03-10 Hao Zhong , Pei Chi , Jiang Zhao , Shenghai Yuan , Xuyang Gao , Thien-Minh Nguyen , Lihua Xie

Panoramic Image Generation (PIG) aims to create coherent images of arbitrary lengths. Most existing methods fall in the joint diffusion paradigm, but their complex and heuristic crop connection designs often limit their ability to achieve…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Teng Zhou , Xiaoyu Zhang , Yongchuan Tang

Video generative models have made remarkable progress, yet they often yield visual artifacts that violate grounding in physical dynamics. Recent works such as PhysGen3D tackle single image-to-3D physics through mesh reconstruction and…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Hwidong Kim , Yunho Kim , Tae-Kyun Kim

Single-image 3D scene reconstruction presents significant challenges due to its inherently ill-posed nature and limited input constraints. Recent advances have explored two promising directions: multiview generative models that train on 3D…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Junlin Hao , Peiheng Wang , Haoyang Wang , Xinggong Zhang , Zongming Guo

360 video object segmentation (360VOS) aims to predict temporally-consistent masks in 360 videos, offering full-scene coverage, benefiting applications, such as VR/AR and embodied AI. Learning 360VOS model is nontrivial due to the lack of…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Dingwen Xiao , Weiming Zhang , Shiqi Wen , Lin Wang

Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Fangfu Liu , Wenqiang Sun , Hanyang Wang , Yikai Wang , Haowen Sun , Junliang Ye , Jun Zhang , Yueqi Duan

In this paper, we tackle the challenging task of Panoramic Image-to-Image translation (Pano-I2I) for the first time. This task is difficult due to the geometric distortion of panoramic images and the lack of a panoramic image dataset with…

计算机视觉与模式识别 · 计算机科学 2023-04-12 Soohyun Kim , Junho Kim , Taekyung Kim , Hwan Heo , Seungryong Kim , Jiyoung Lee , Jin-Hwa Kim

In this paper we propose a geometry-aware model for video object detection. Specifically, we consider the setting that cameras can be well approximated as static, e.g. in video surveillance scenarios, and scene pseudo depth maps can…

计算机视觉与模式识别 · 计算机科学 2019-09-10 Dan Xu , Weidi Xie , Andrew Zisserman
‹ 上一页 1 8 9 10 下一页 ›