中文
相关论文

相关论文: Style4D-Bench: A Benchmark Suite for 4D Stylizatio…

200 篇论文

High-quality 4D reconstruction enables photorealistic and immersive rendering of the dynamic real world. However, unlike static scenes that can be fully captured with a single camera, high-quality dynamic scenes typically require dense…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Weihong Pan , Xiaoyu Zhang , Zhuang Zhang , Zhichao Ye , Nan Wang , Haomin Liu , Guofeng Zhang

4D spatial intelligence involves perceiving and processing how objects move or change over time. Humans naturally possess 4D spatial intelligence, supporting a broad spectrum of spatial reasoning abilities. To what extent can Multimodal…

We introduce Control4D, an innovative framework for editing dynamic 4D portraits using text instructions. Our method addresses the prevalent challenges in 4D editing, notably the inefficiencies of existing 4D representations and the…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Ruizhi Shao , Jingxiang Sun , Cheng Peng , Zerong Zheng , Boyao Zhou , Hongwen Zhang , Yebin Liu

We introduce G-Style, a novel algorithm designed to transfer the style of an image onto a 3D scene represented using Gaussian Splatting. Gaussian Splatting is a powerful 3D representation for novel view synthesis, as -- compared to other…

图形学 · 计算机科学 2024-09-06 Áron Samuel Kovács , Pedro Hermosilla , Renata G. Raidou

We present Stylos, a single-forward 3D Gaussian framework for 3D style transfer that operates on unposed content, from a single image to a multi-view collection, conditioned on a separate reference style image. Stylos synthesizes a stylized…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Hanzhou Liu , Jia Huang , Mi Lu , Srikanth Saripalli , Peng Jiang

Previous text-to-4D methods have leveraged multiple Score Distillation Sampling (SDS) techniques, combining motion priors from video-based diffusion models (DMs) with geometric priors from multiview DMs to implicitly guide 4D renderings.…

计算机视觉与模式识别 · 计算机科学 2024-11-20 Qiaowei Miao , JinSheng Quan , Kehan Li , Yawei Luo

Recent progress in pre-trained diffusion models and 3D generation have spurred interest in 4D content creation. However, achieving high-fidelity 4D generation with spatial-temporal consistency remains a challenge. In this work, we propose…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Yifei Zeng , Yanqin Jiang , Siyu Zhu , Yuanxun Lu , Youtian Lin , Hao Zhu , Weiming Hu , Xun Cao , Yao Yao

Layout-guided text-to-image models offer greater control over the generation process by explicitly conditioning image synthesis on the spatial arrangement of elements. As a result, their adoption has increased in many computer vision…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Elena Izzo , Luca Parolari , Davide Vezzaro , Lamberto Ballan

3D style transfer refers to the artistic stylization of 3D assets based on reference style images. Recently, 3DGS-based stylization methods have drawn considerable attention, primarily due to their markedly enhanced training and rendering…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Yian Zhao , Rushi Ye , Ruochong Zheng , Zesen Cheng , Chaoran Feng , Jiashu Yang , Pengchong Qiao , Chang Liu , Jie Chen

In recent years, there has been a growing demand to stylize a given 3D scene to align with the artistic style of reference images for creative purposes. While 3D Gaussian Splatting(GS) has emerged as a promising and efficient method for…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Yangkai Lin , Jiabao Lei , Kui jia

We introduce ReStyle3D, a novel framework for scene-level appearance transfer from a single style image to a real-world scene represented by multiple views. The method combines explicit semantic correspondences with multi-view consistency…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Liyuan Zhu , Shengqu Cai , Shengyu Huang , Gordon Wetzstein , Naji Khosravan , Iro Armeni

Gaussian Splatting has been considered as a novel way for view synthesis of dynamic scenes, which shows great potential in AIoT applications such as digital twins. However, recent dynamic Gaussian Splatting methods significantly degrade…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Yiwei Li , Jiannong Cao , Penghui Ruan , Divya Saxena , Songye Zhu , Yinfeng Cao

Recent 4D dynamic scene editing methods require editing thousands of 2D images used for dynamic scene synthesis and updating the entire scene with additional training loops, resulting in several hours of processing to edit a single dynamic…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Joohyun Kwon , Hanbyel Cho , Junmo Kim

Indoor environments evolve as objects move, appear, or leave the scene. Capturing these dynamics requires maintaining temporally consistent instance identities across intermittently captured 3D scans, even when changes are unobserved. We…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Emily Steiner , Jianhao Zheng , Henry Howard-Jenkins , Chris Xie , Iro Armeni

Creating large-scale virtual urban scenes with variant styles is inherently challenging. To facilitate prototypes of virtual production and bypass the need for complex materials and lighting setups, we introduce the first…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Yingshu Chen , Huajian Huang , Tuan-Anh Vu , Ka Chun Shum , Sai-Kit Yeung

Persistent dynamic scene modeling for tracking and novel-view synthesis remains challenging due to the difficulty of capturing accurate deformations while maintaining computational efficiency. We propose SCas4D, a cascaded optimization…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Jipeng Lyu , Jiahua Dong , Yu-Xiong Wang

Dynamic scene rendering opens new avenues in autonomous driving by enabling closed-loop simulations with photorealistic data, which is crucial for validating end-to-end algorithms. However, the complex and highly dynamic nature of traffic…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Rui Song , Chenwei Liang , Yan Xia , Walter Zimmer , Hu Cao , Holger Caesar , Andreas Festag , Alois Knoll

Reconstructing dynamic 3D scenes from 2D images and generating diverse views over time is challenging due to scene complexity and temporal dynamics. Despite advancements in neural implicit models, limitations persist: (i) Inadequate Scene…

计算机视觉与模式识别 · 计算机科学 2024-02-23 Zeyu Yang , Hongye Yang , Zijie Pan , Li Zhang

Text-to-3D (T23D) generation has emerged as a crucial visual generation task, aiming at synthesizing 3D content from textual descriptions. Studies of this task are currently shifting from per-scene T23D, which requires optimization of the…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Xiao Cai , Sitong Su , Jingkuan Song , Pengpeng Zeng , Ji Zhang , Qinhong Du , Mengqi Li , Heng Tao Shen , Lianli Gao

Generating high-quality 4D content from monocular videos for applications such as digital humans and AR/VR poses challenges in ensuring temporal and spatial consistency, preserving intricate details, and incorporating user guidance…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Minghao Yin , Yukang Cao , Songyou Peng , Kai Han