中文
相关论文

相关论文: Versatile Video Tokenization with Generative 2D Ga…

200 篇论文

Recent advances in 3D Gaussian Splatting have allowed for real-time, high-fidelity novel view synthesis. Nonetheless, these models have significant storage requirements for large and medium-sized scenes, hindering their deployment over…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Diego Revilla , Pooja Suresh , Anand Bhojan , Ooi Wei Tsang

3D Gaussian Splatting (3DGS) has established itself as an efficient representation for real-time, high-fidelity 3D scene reconstruction. However, scaling 3DGS to large and unbounded scenes such as city blocks remains difficult. Existing…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Sheng-Hsiang Hung , Ting-Yu Yen , Wei-Fang Sun , Simon See , Shih-Hsuan Hung , Hung-Kuo Chu

This work presents VTok, a unified video tokenization framework that can be used for both generation and understanding tasks. Unlike the leading vision-language systems that tokenize videos through a naive frame-sampling strategy, we…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Feng Wang , Yichun Shi , Ceyuan Yang , Qiushan Guo , Jingxiang Sun , Alan Yuille , Peng Wang

Novel view synthesis (NVS) of static and dynamic urban scenes is essential for autonomous driving simulation, yet existing methods often struggle to balance reconstruction time with quality. While state-of-the-art neural radiance fields and…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Sheng Miao , Sijin Li , Pan Wang , Dongfeng Bai , Bingbing Liu , Yue Wang , Andreas Geiger , Yiyi Liao

Gaussian Splatting (GS) is a popular approach for 3D reconstruction, mostly due to its ability to converge reasonably fast, faithfully represent the scene and render (novel) views in a fast fashion. However, it suffers from large storage…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Anil Armagan , Albert Saà-Garriga , Bruno Manganelli , Kyuwon Kim , M. Kerim Yucel

Recently, high-fidelity scene reconstruction with an optimized 3D Gaussian splat representation has been introduced for novel view synthesis from sparse image sets. Making such representations suitable for applications like network…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Simon Niedermayr , Josef Stumpfegger , Rüdiger Westermann

Video tokenizers are essential for latent video diffusion models, converting raw video data into spatiotemporally compressed latent spaces for efficient training. However, extending state-of-the-art video tokenizers to achieve a temporal…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Aniruddha Mahapatra , Long Mai , David Bourgin , Yitian Zhang , Feng Liu

3D Gaussian Splatting (3DGS) enables real-time novel view synthesis with high visual quality. However, existing methods struggle with semi-transparent specular surfaces that exhibit both complex reflections and clear transmission, often…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Ji Shi , Xianghua Ying , Bowei Xing , Ruohao Guo , Wenzhen Yue

Traditional video summarization methods generate fixed video representations regardless of user interest. Therefore such methods limit users' expectations in content search and exploration scenarios. Multi-modal video summarization is one…

计算机视觉与模式识别 · 计算机科学 2021-04-27 Jia-Hong Huang , Luka Murn , Marta Mrak , Marcel Worring

Recently, reconstructing scenes from a single panoramic image using advanced 3D Gaussian Splatting (3DGS) techniques has attracted growing interest. Panoramic images offer a 360$\times$ 180 field of view (FoV), capturing the entire scene in…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Zhijie Shen , Chunyu Lin , Shujuan Huang , Lang Nie , Kang Liao , Yao Zhao

The usage of deep generative models for image compression has led to impressive performance gains over classical codecs while neural video compression is still in its infancy. Here, we propose an end-to-end, deep generative modeling…

计算机视觉与模式识别 · 计算机科学 2019-11-05 Jun Han , Salvator Lombardo , Christopher Schroers , Stephan Mandt

Current 4D Gaussian frameworks for dynamic scene reconstruction deliver impressive visual fidelity and rendering speed, however, the inherent trade-off between storage costs and the ability to characterize complex physical motions…

图形学 · 计算机科学 2025-07-11 Wei Yao , Shuzhao Xie , Letian Li , Weixiang Zhang , Zhixin Lai , Shiqi Dai , Ke Zhang , Zhi Wang

Novel view synthesis of dynamic scenes is becoming important in various applications, including augmented and virtual reality. We propose a novel 4D Gaussian Splatting (4DGS) algorithm for dynamic scenes from casually recorded monocular…

计算机视觉与模式识别 · 计算机科学 2024-11-14 Mijeong Kim , Jongwoo Lim , Bohyung Han

Jointly estimating camera poses and mapping scenes from RGBD images is a fundamental task in simultaneous localization and mapping (SLAM). State-of-the-art methods employ 3D Gaussians to represent a scene, and render these Gaussians through…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Pengchong Hu , Zhizhong Han

Recovering 3D information from scenes via multi-view stereo reconstruction (MVS) and novel view synthesis (NVS) is inherently challenging, particularly in scenarios involving sparse-view setups. The advent of 3D Gaussian Splatting (3DGS)…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Shubhendu Jena , Shishir Reddy Vutukur , Adnane Boukhayma

Feed-forward 3D Gaussian Splatting methods enable single-pass reconstruction and real-time rendering. However, they typically adopt rigid pixel-to-Gaussian or voxel-to-Gaussian pipelines that uniformly allocate Gaussians, leading to…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Injae Kim , Chaehyeon Kim , Minseong Bae , Minseok Joo , Hyunwoo J. Kim

Reconstructing dynamic scenes from video sequences is a highly promising task in the multimedia domain. While previous methods have made progress, they often struggle with slow rendering and managing temporal complexities such as…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Jinbo Yan , Rui Peng , Luyang Tang , Ronggang Wang

Reconstructing urban scenes is challenging due to their complex geometries and the presence of potentially dynamic objects. 3D Gaussian Splatting (3DGS)-based methods have shown strong performance, but existing approaches often incorporate…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Ziwen Li , Jiaxin Huang , Runnan Chen , Yunlong Che , Yandong Guo , Tongliang Liu , Fakhri Karray , Mingming Gong

Building Free-Viewpoint Videos in a streaming manner offers the advantage of rapid responsiveness compared to offline training methods, greatly enhancing user experience. However, current streaming approaches face challenges of high…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Jinbo Yan , Rui Peng , Zhiyan Wang , Luyang Tang , Jiayu Yang , Jie Liang , Jiahao Wu , Ronggang Wang

Learned video compression (LVC) has witnessed remarkable advancements in recent years. Similar as the traditional video coding, LVC inherits motion estimation/compensation, residual coding and other modules, all of which are implemented…

图像与视频处理 · 电气工程与系统科学 2023-09-22 Yanbo Gao , Wenjia Huang , Shuai Li , Hui Yuan , Mao Ye , Siwei Ma