中文
相关论文

相关论文: PanoWorld: Geometry-Consistent Panoramic Video Wor…

200 篇论文

Synthesis and reconstruction of 3D human head has gained increasing interests in computer vision and computer graphics recently. Existing state-of-the-art 3D generative adversarial networks (GANs) for 3D human head synthesis are either…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Sizhe An , Hongyi Xu , Yichun Shi , Guoxian Song , Umit Ogras , Linjie Luo

The rapid advancement of diffusion models holds the promise of revolutionizing the application of VR and AR technologies, which typically require scene-level 4D assets for user experience. Nonetheless, existing diffusion models…

计算机视觉与模式识别 · 计算机科学 2025-05-14 Haiyang Zhou , Wangbo Yu , Jiawen Guan , Xinhua Cheng , Yonghong Tian , Li Yuan

Maintaining spatial world consistency over long horizons remains a central challenge for camera-controllable video generation. Existing memory-based approaches often condition generation on globally reconstructed 3D scenes by rendering…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Zun Wang , Han Lin , Jaehong Yoon , Jaemin Cho , Yue Zhang , Mohit Bansal

Recent video generators achieve striking photorealism, yet remain fundamentally inconsistent in 3D. We present WorldReel, a 4D video generator that is natively spatio-temporally consistent. WorldReel jointly produces RGB frames together…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Shaoheng Fang , Hanwen Jiang , Yunpeng Bai , Niloy J. Mitra , Qixing Huang

Achieving an immersive experience enabling users to explore virtual environments with six degrees of freedom (6DoF) is essential for various applications such as virtual reality (VR). Wide-baseline panoramas are commonly used in these…

计算机视觉与模式识别 · 计算机科学 2023-12-07 Zheng Chen , Yan-Pei Cao , Yuan-Chen Guo , Chen Wang , Ying Shan , Song-Hai Zhang

Being able to edit panoramic images is crucial for creating realistic 360{\deg} visual experiences. However, existing perspective-based image editing methods fail to model the spatial structure of panoramas. Conventional cube-map…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Dong Liang , Yuhao Liu , Jinyuan Jia , Youjun Zhao , Rynson W. H. Lau

Generative video models are increasingly studied as implicit world models, yet evaluating whether they produce physically plausible 3D structure and motion remains challenging. Most existing video evaluation pipelines rely heavily on human…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Jiaxin Wu , Yihao Pi , Yinling Zhang , Yuheng Li , Xueyan Zou

We present a novel video generation framework that integrates 3-dimensional geometry and dynamic awareness. To achieve this, we augment 2D videos with 3D point trajectories and align them in pixel space. The resulting 3D-aware video…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Yunuo Chen , Junli Cao , Vidit Goel , Sergei Korolev , Chenfanfu Jiang , Jian Ren , Sergey Tulyakov , Anil Kag

3D Visual Grounding (3DVG) is a critical bridge from vision-language perception to robotics, requiring both language understanding and 3D scene reasoning. Traditional supervised models leverage explicit 3D geometry but exhibit limited…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Seongmin Jung , Seongho Choi , Gunwoo Jeon , Minsu Cho , Jongwoo Lim

Current compositional image-to-3D scene generation approaches construct 3D scenes by time-consuming iterative layout optimization or inflexible joint object-layout generation. Moreover, most methods rely on limited field-of-view perspective…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Zidian Qiu , Ancong Wu

Explorable 3D world generation from a single image or text prompt forms a cornerstone of spatial intelligence. Recent works utilize video model to achieve wide-scope and generalizable 3D world generation. However, existing approaches often…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Zhongqi Yang , Wenhang Ge , Yuqi Li , Jiaqi Chen , Haoyuan Li , Mengyin An , Fei Kang , Hua Xue , Baixin Xu , Yuyang Yin , Eric Li , Yang Liu , Yikai Wang , Hao-Xiang Guo , Yahui Zhou

Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Gwanghyun Kim , Xueting Li , Ye Yuan , Koki Nagano , Tianye Li , Jan Kautz , Se Young Chun , Umar Iqbal

Real-world applications like video gaming and virtual reality often demand the ability to model 3D scenes that users can explore along custom camera trajectories. While significant progress has been made in generating 3D objects from text…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Tianyu Huang , Wangguandong Zheng , Tengfei Wang , Yuhao Liu , Zhenwei Wang , Junta Wu , Jie Jiang , Hui Li , Rynson W. H. Lau , Wangmeng Zuo , Chunchao Guo

Video diffusion models generate high-quality and diverse worlds; however, individual frames often lack 3D consistency across the output sequence, which makes the reconstruction of 3D worlds difficult. To this end, we propose a new method…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Lukas Höllein , Matthias Nießner

State-of-the-art video generation models produce remarkable photorealism, but they lack the precise control required to align generated content with specific scene requirements. Furthermore, without an underlying explicit geometry, these…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dana Cohen-Bar , Ido Sobol , Raphael Bensadoun , Shelly Sheynin , Oran Gafni , Or Patashnik , Daniel Cohen-Or , Amit Zohar

Panoramic video generation has attracted growing attention due to its applications in virtual reality and immersive media. However, existing methods lack explicit motion control and struggle to generate scenes with large and complex…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Cheng Zhang , Hanwen Liang , Donny Y. Chen , Qianyi Wu , Konstantinos N. Plataniotis , Camilo Cruz Gambardella , Jianfei Cai

Immersive scene generation, notably panorama creation, benefits significantly from the adaptation of large pre-trained text-to-image (T2I) models for multi-view image generation. Due to the high cost of acquiring multi-view images,…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Aoming Liu , Zhong Li , Zhang Chen , Nannan Li , Yi Xu , Bryan A. Plummer

Generating long-range, geometrically consistent video presents a fundamental dilemma: while consistency demands strict adherence to 3D geometry in pixel space, state-of-the-art generative models operate most effectively in a…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Hanyang Kong , Xingyi Yang , Xiaoxu Zheng , Xinchao Wang

We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Yang-Tian Sun , Zehuan Huang , Yifan Niu , Lin Ma , Yan-Pei Cao , Yuewen Ma , Xiaojuan Qi

Video Motion Magnification (VMM) reveals imperceptible dynamics but often suffers from structural inconsistencies under complex geometric transformations. Existing learning-based methods generally face a trade-off between the limited global…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Kecheng Han , Yuchen Zhang , Bingqing Liu , Boqiang Guo , Wenbin Zheng , Shiyuan Pei