中文
相关论文

相关论文: Modeling and Driving Human Body Soundfields throug…

200 篇论文

This work presents Interactive Conversational 3D Virtual Human (ICo3D), a method for generating an interactive, conversational, and photorealistic 3D human avatar. Based on multi-view captures of a subject, we create an animatable 3D face…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Richard Shaw , Youngkyoon Jang , Athanasios Papaioannou , Arthur Moreau , Helisa Dhamo , Zhensong Zhang , Eduardo Pérez-Pellitero

We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try…

图形学 · 计算机科学 2024-09-23 Matthew Caren , Kartik Chandra , Joshua B. Tenenbaum , Jonathan Ragan-Kelley , Karima Ma

Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this work, we challenge this paradigm with WavFlow, a framework that generates high-fidelity…

Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-supervised audio…

声音 · 计算机科学 2019-05-15 Yu-Ding Lu , Hsin-Ying Lee , Hung-Yu Tseng , Ming-Hsuan Yang

Prevailing Video-to-Audio (V2A) generation models operate offline, assuming an entire video sequence or chunks of frames are available beforehand. This critically limits their use in interactive applications such as live content creation…

In this paper, we introduce GaussianMotion, a novel human rendering model that generates fully animatable scenes aligned with textual descriptions using Gaussian Splatting. Although existing methods achieve reasonable text-to-3D generation…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Gyumin Shim , Sangmin Lee , Jaegul Choo

Accurate 3D reconstruction of vertebral anatomy from ultrasound is important for guiding minimally invasive spine interventions, but it remains challenging due to acoustic shadowing and view-dependent signal variations. We propose an…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Magdalena Wysocki , Kadir Burak Buldu , Miruna-Alexandra Gafencu , Mohammad Farid Azampour , Nassir Navab

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between current results and…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Xiaozhong Ji , Chuming Lin , Zhonggan Ding , Ying Tai , Junwei Zhu , Xiaobin Hu , Donghao Luo , Yanhao Ge , Chengjie Wang

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

计算机视觉与模式识别 · 计算机科学 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

Non-interactive and linear experiences like cinema film offer high quality surround sound audio to enhance immersion, however the listener's experience is usually fixed to a single acoustic perspective. With the rise of virtual reality,…

音频与语音处理 · 电气工程与系统科学 2024-10-30 Lachlan Birnie , Thushara Abhayapala , Vladimir Tourbabin , Prasanga Samarasinghe

Ambisonics is a method for capturing and rendering a sound field accurately, assuming that the acoustics of the playback room does not significantly influence the sound field. However, in practice, the acoustics of the playback room may…

音频与语音处理 · 电气工程与系统科学 2025-10-14 Ali Fallah , Shun Nakamura , Steven van de Par

Speech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but slow or fast yet temporally unstable. Diffusion methods…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Madhav Agarwal , Mingtian Zhang , Laura Sevilla-Lara , Steven McDonagh

Developing algorithms for sound classification, detection, and localization requires large amounts of flexible and realistic audio data, especially when leveraging modern machine learning and beamforming techniques. However, most existing…

音频与语音处理 · 电气工程与系统科学 2026-01-23 Luca Barbisan , Marco Levorato , Fabrizio Riente

Artificial neural networks are increasingly powerful models of brain computation, yet it remains unclear whether improving their performance in downstream tasks also makes their internal representations more similar to brain signals. To…

机器学习 · 计算机科学 2026-03-05 Leonardo Pepino , Pablo Riera , Juan Kamienkowski , Luciana Ferrer

We present a novel hybrid sound propagation algorithm for interactive applications. Our approach is designed for dynamic scenes and uses a neural network-based learned scattered field representation along with ray tracing to generate…

声音 · 计算机科学 2021-09-28 Zhenyu Tang , Hsien-Yu Meng , Dinesh Manocha

We present a novel framework for animating humans in 3D scenes using 3D Gaussian Splatting (3DGS), a neural scene representation that has recently achieved state-of-the-art photorealistic results for novel-view synthesis but remains…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Aymen Mir , Jian Wang , Riza Alp Guler , Chuan Guo , Gerard Pons-Moll , Bing Zhou

Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly, they often struggle under visually degraded conditions such as motion blur or occlusions. In…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Daniel Adebi , Sagnik Majumder , Kristen Grauman

Implicit Neural Representations (INRs) are nowadays used to represent multimedia signals across various real-life applications, including image super-resolution, image compression, or 3D rendering. Existing methods that leverage INRs are…

机器学习 · 计算机科学 2023-06-21 Filip Szatkowski , Karol J. Piczak , Przemysław Spurek , Jacek Tabor , Tomasz Trzciński

Dynamic 3D scene representation and novel view synthesis are crucial for enabling immersive experiences required by AR/VR and metaverse applications. It is a challenging task due to the complexity of unconstrained real-world scenes and…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Zeyu Yang , Zijie Pan , Xiatian Zhu , Li Zhang , Jianfeng Feng , Yu-Gang Jiang , Philip H. S. Torr

Human-like environment recognition by musculoskeletal humanoids is important for task realization in real complex environments and for use as dummies for test subjects. Humans integrate various sensory information to perceive their…