English
Related papers

Related papers: Audio-Plane: Audio Factorization Plane Gaussian Sp…

200 papers

Differentiable rendering techniques have recently shown promising results for free-viewpoint video synthesis of characters. However, such methods, either Gaussian Splatting or neural implicit rendering, typically necessitate per-subject…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Boyao Zhou , Shunyuan Zheng , Hanzhang Tu , Ruizhi Shao , Boning Liu , Shengping Zhang , Liqiang Nie , Yebin Liu

Utterance clustering is one of the actively researched topics in audio signal processing and machine learning. This study aims to improve the performance of utterance clustering by processing multichannel (stereo) audio signals. Processed…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-22 Yingjun Dong , Neil G. MacLaren , Yiding Cao , Francis J. Yammarino , Shelley D. Dionne , Michael D. Mumford , Shane Connelly , Hiroki Sayama , Gregory A. Ruark

Talking head synthesis, an advanced method for generating portrait videos from a still image driven by specific content, has garnered widespread attention in virtual reality, augmented reality and game production. Recently, significant…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Ming Meng , Yufei Zhao , Bo Zhang , Yonggui Zhu , Weimin Shi , Maxwell Wen , Zhaoxin Fan

3D occupancy prediction is critical for comprehensive scene understanding in vision-centric autonomous driving. Recent advances have explored utilizing 3D semantic Gaussians to model occupancy while reducing computational overhead, but they…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Xiaoyang Yan , Muleilan Pei , Shaojie Shen

Empowering 3D Gaussian Splatting with generalization ability is appealing. However, existing generalizable 3D Gaussian Splatting methods are largely confined to narrow-range interpolation between stereo images due to their heavy backbones,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Yunsong Wang , Tianxin Huang , Hanlin Chen , Gim Hee Lee

High-fidelity 3D Gaussian Splatting methods excel at capturing fine textures but often overlook model compactness, resulting in massive splat counts, bloated memory, long training, and complex post-processing. We present Micro-Splatting:…

Graphics · Computer Science 2025-09-03 Jee Won Lee , Hansol Lim , Sooyeun Yang , Jongseong Brad Choi

Audio domain transfer is the process of modifying audio signals to match characteristics of a different domain, while retaining the original content. This paper investigates the potential of Gaussian Flow Bridges, an emerging approach in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-31 Eloi Moliner , Sebastian Braun , Hannes Gamper

Modern scene reconstruction methods, such as 3D Gaussian Splatting, deliver photo-realistic novel view synthesis at real-time speeds, yet their adoption in interactive graphics applications has been limited. A major bottleneck is the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Amr Sharafeldin , Shrisudhan Govindarajan , Thomas Walker , Aryan Mikaeili , Daniel Rebain , Kwang Moo Yi , Andrea Tagliasacchi

This paper is about developing personalized speech synthesis systems with recordings of mildly impaired speech. In particular, we consider consonant and vowel alterations resulted from partial glossectomy, the surgical removal of part of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-10 Yusheng Tian , Guangyan Zhang , Tan Lee

Latent variable generative models have emerged as powerful tools for generative tasks including image and video synthesis. These models are enabled by pretrained autoencoders that map high resolution data into a compressed lower dimensional…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Mohammed Suhail , Carlos Esteves , Leonid Sigal , Ameesh Makadia

Ultrasound imaging is a cornerstone of non-invasive clinical diagnostics, yet its limited field of view poses challenges for novel view synthesis. We present UltraGS, a real-time framework that adapts Gaussian Splatting to sensorless…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Yuezhe Yang , Qingqing Ruan , Wenjie Cai , Yudang Dong , Dexin Yang , Xingbo Dong , Zhe Jin , Yong Dai

Reconstructing dynamic 3D scenes from 2D images and generating diverse views over time is challenging due to scene complexity and temporal dynamics. Despite advancements in neural implicit models, limitations persist: (i) Inadequate Scene…

Computer Vision and Pattern Recognition · Computer Science 2024-02-23 Zeyu Yang , Hongye Yang , Zijie Pan , Li Zhang

Recent advancements in speech-driven 3D talking head generation have made significant progress in lip synchronization. However, existing models still struggle to capture the perceptual alignment between varying speech characteristics and…

Graphics · Computer Science 2025-04-01 Lee Chae-Yeon , Oh Hyun-Bin , Han EunGi , Kim Sung-Bin , Suekyeong Nam , Tae-Hyun Oh

Given an arbitrary audio clip, audio-driven 3D facial animation aims to generate lifelike lip motions and facial expressions for a 3D head. Existing methods typically rely on training their models using limited public 3D datasets that…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Liying Lu , Tianke Zhang , Yunfei Liu , Xuangeng Chu , Yu Li

Audio-driven talking head generation is advancing from 2D to 3D content. Notably, Neural Radiance Field (NeRF) is in the spotlight as a means to synthesize high-quality 3D talking head outputs. Unfortunately, this NeRF-based approach…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Gihoon Kim , Kwanggyoon Seo , Sihun Cha , Junyong Noh

Recently, 3D GANs based on 3D Gaussian splatting have been proposed for high quality synthesis of human heads. However, existing methods stabilize training and enhance rendering quality from steep viewpoints by conditioning the random…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Florian Barthel , Wieland Morgenstern , Paul Hinzer , Anna Hilsmann , Peter Eisert

Although automatically animating audio-driven talking heads has recently received growing interest, previous efforts have mainly concentrated on achieving lip synchronization with the audio, neglecting two crucial elements for generating…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Shuai Tan , Bin Ji , Ye Pan

3D Gaussian Splatting has recently gained traction for its efficient training and real-time rendering. While its vanilla representation is mainly designed for view synthesis, recent works extended it to scene understanding with language…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Siyun Liang , Sen Wang , Kunyi Li , Michael Niemeyer , Stefano Gasperini , Hendrik P. A. Lensch , Nassir Navab , Federico Tombari

Audio fingerprinting techniques have seen great advances in recent years, enabling accurate and fast audio retrieval even in conditions when the queried audio sample has been highly deteriorated or recorded in noisy conditions. Expectedly,…

Information Retrieval · Computer Science 2025-09-26 Kemal Altwlkany , Sead Delalić , Adis Alihodžić , Elmedin Selmanović , Damir Hasić

Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computational costs. Some…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Ziqi Ni , Ao Fu , Yi Zhou
‹ Prev 1 8 9 10 Next ›