中文
相关论文

相关论文: Versatile Video Tokenization with Generative 2D Ga…

200 篇论文

Implicit Neural Representations (INRs) employ neural networks to approximate discrete data as continuous functions. In the context of video data, such models can be utilized to transform the coordinates of pixel locations along with frame…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Weronika Smolak-Dyżewska , Dawid Malarz , Kornel Howil , Jan Kaczmarczyk , Marcin Mazur , Przemysław Spurek

The accurate reconstruction of dynamic street scenes is critical for applications in autonomous driving, augmented reality, and virtual reality. Traditional methods relying on dense point clouds and triangular meshes struggle with moving…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Peizhen Zheng , Dongjing Jiang , Qingchong Jiao , Redouane EL Bouchtaoui , Flynnwell Jianfei Zhang

3D Gaussian Splatting (GS) has emerged as a powerful representation for high-quality scene reconstruction, offering compelling rendering quality. However, the training process of GS often suffers from slow convergence due to inefficient…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Binxiao Huang , Zhengwu Liu , Ngai Wong

Novel view synthesis of dynamic scenes has been an intriguing yet challenging problem. Despite recent advancements, simultaneously achieving high-resolution photorealistic results, real-time rendering, and compact storage remains a…

计算机视觉与模式识别 · 计算机科学 2024-04-08 Zhan Li , Zhang Chen , Zhong Li , Yi Xu

We propose a method to enhance 3D Gaussian Splatting (3DGS)~\cite{Kerbl2023}, addressing challenges in initialization, optimization, and density control. Gaussian Splatting is an alternative for rendering realistic images while supporting…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Xingjun Wang , Lianlei Shan

Leveraging text, images, structure maps, or motion trajectories as conditional guidance, diffusion models have achieved great success in automated and high-quality video generation. However, generating smooth and rational transition videos…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Zuhao Yang , Jiahui Zhang , Yingchen Yu , Shijian Lu , Song Bai

Although 3D Gaussian Splatting (3D-GS) achieves efficient rendering for novel view synthesis, extending it to dynamic scenes still results in substantial memory overhead from replicating Gaussians across frames. To address this challenge,…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Chun-Tin Wu , Jun-Cheng Chen

Tensor singular value decomposition (t-SVD) is a promising tool for multi-dimensional image representation, which decomposes a multi-dimensional image into a latent tensor and an accompanying transform matrix. However, two critical…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Yiming Zeng , Xi-Le Zhao , Wei-Hao Wu , Teng-Yu Ji , Chao Wang

We introduce GS2E (Gaussian Splatting to Event), a large-scale synthetic event dataset for high-fidelity event vision tasks, captured from real-world sparse multi-view RGB images. Existing event datasets are often synthesized from dense RGB…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Yuchen Li , Chaoran Feng , Zhenyu Tang , Kaiyuan Deng , Wangbo Yu , Yonghong Tian , Li Yuan

3D Gaussian Splatting (3DGS) has emerged as a powerful representation due to its efficiency and high-fidelity rendering. 3DGS training requires a known camera pose for each input view, typically obtained by Structure-from-Motion (SfM)…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Zhen-Hui Dong , Sheng Ye , Yu-Hui Wen , Nannan Li , Yong-Jin Liu

Implicit neural representations for video have been recognized as a novel and promising form of video representation. Existing works pay more attention to improving video reconstruction quality but little attention to the decoding speed.…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Zhizhuo Pang , Zhihui Ke , Xiaobo Zhou , Tie Qiu

We present a framework that enables fast reconstruction and real-time rendering of urban-scale scenes while maintaining robustness against appearance variations across multi-view captures. Our approach begins with scene partitioning for…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Zhensheng Yuan , Haozhi Huang , Zhen Xiong , Di Wang , Guanghua Yang

We propose HoliGS, a novel deformable Gaussian splatting framework that addresses embodied view synthesis from long monocular RGB videos. Unlike prior 4D Gaussian splatting and dynamic NeRF pipelines, which struggle with training overhead…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Xiaoyuan Wang , Yizhou Zhao , Botao Ye , Xiaojun Shan , Weijie Lyu , Lu Qi , Kelvin C. K. Chan , Yinxiao Li , Ming-Hsuan Yang

Recently, 3D Gaussian Splatting (3DGS) has revolutionized radiance field reconstruction, manifesting efficient and high-fidelity novel view synthesis. However, accurately representing surfaces, especially in large and complex scenarios,…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Yang Liu , Chuanchen Luo , Zhongkai Mao , Junran Peng , Zhaoxiang Zhang

Modeling dynamic, large-scale urban scenes is challenging due to their highly intricate geometric structures and unconstrained dynamics in both space and time. Prior methods often employ high-level architectural priors, separating static…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Yurui Chen , Chun Gu , Junzhe Jiang , Xiatian Zhu , Li Zhang

Differentiable rendering techniques have recently shown promising results for free-viewpoint video synthesis of characters. However, such methods, either Gaussian Splatting or neural implicit rendering, typically necessitate per-subject…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Boyao Zhou , Shunyuan Zheng , Hanzhang Tu , Ruizhi Shao , Boning Liu , Shengping Zhang , Liqiang Nie , Yebin Liu

3D reconstruction from multi-view images is a core challenge in computer vision. Recently, feed-forward methods have emerged as efficient and robust alternatives to traditional per-scene optimization techniques. Among them, state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Zipeng Wang , Dan Xu

Reconstructing static 3D scene from monocular video with dynamic objects is important for numerous applications such as virtual reality and autonomous driving. Current approaches typically rely on background for static scene reconstruction,…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Yedong Shen , Shiqi Zhang , Sha Zhang , Yifan Duan , Xinran Zhang , Wenhao Yu , Lu Zhang , Jiajun Deng , Yanyong Zhang

In this work, we revisit several key design choices of modern Transformer-based approaches for feed-forward 3D Gaussian Splatting (3DGS) prediction. We argue that the common practice of regressing Gaussian means as depths along camera rays…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Jiawei Ren , Michal Jan Tyszkiewicz , Jiahui Huang , Zan Gojcic

While the generation of 3D content from single-view images has been extensively studied, the creation of physically consistent 3D dynamic scenes from videos remains in its early stages. We propose a novel framework leveraging generative 3D…

图形学 · 计算机科学 2025-10-31 Zhiwei Zhao , Alan Zhao , Minchen Li , Yixin Hu