中文
相关论文

相关论文: Ambisonizer: Neural Upmixing as Spherical Harmonic…

200 篇论文

Spatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic encoding methods often rely on spherical microphone arrays for…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Yue Qiao , Vinay Kothapally , Meng Yu , Dong Yu

Ambisonics is a complete theory for spatial audio whose building blocks are the spherical harmonics. Some of the drawbacks of low order Ambisonics, like poor source directivity and small sweet-spot, are directly related to the properties of…

声音 · 计算机科学 2020-03-09 Davide Scaini

Ambisonics is a method for capturing and rendering a sound field accurately, assuming that the acoustics of the playback room does not significantly influence the sound field. However, in practice, the acoustics of the playback room may…

音频与语音处理 · 电气工程与系统科学 2025-10-14 Ali Fallah , Shun Nakamura , Steven van de Par

In the rapidly evolving fields of virtual and augmented reality, accurate spatial audio capture and reproduction are essential. For these applications, Ambisonics has emerged as a standard format. However, existing methods for encoding…

音频与语音处理 · 电气工程与系统科学 2024-11-27 Yhonatan Gayer , Vladimir Tourbabin , Zamir Ben-Hur , Jacob Donley , Boaz Rafaely

Ambisonics encoding of microphone array signals can enable various spatial audio applications, such as virtual reality or telepresence, but it is typically designed for uniformly-spaced spherical microphone arrays. This paper proposes a…

音频与语音处理 · 电气工程与系统科学 2024-01-12 Mikko Heikkinen , Archontis Politis , Tuomas Virtanen

Ambisonics is a scene-based spatial audio format that has several useful features compared to object-based formats, such as efficient whole scene rotation and versatility. However, it does not provide direct access to the individual source…

声音 · 计算机科学 2023-06-21 Francesc Lluís , Nils Meyer-Kahlen , Vasileios Chatziioannou , Alex Hofmann

Ambisonics, a popular format of spatial audio, is the spherical harmonic (SH) representation of the plane wave density function of a sound field. Many algorithms operate in the SH domain and utilize the Ambisonics as their input signal. The…

音频与语音处理 · 电气工程与系统科学 2024-03-01 Bar Shaybet , Anurag Kumar , Vladimir Tourbabin , Boaz Rafaely

Generating a stereophonic presentation from a monophonic audio signal is a challenging open task, especially if the goal is to obtain a realistic spatial imaging with a specific panning of sound elements. In this work, we propose to convert…

Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Existing audio-video generation models often fail to maintain…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Shihao Cheng , Jiaxu Zhang , Quanyue Song , Shansong Liu , Zhizhi Guo , Xiaolei Zhang , Chi Zhang , Xuelong Li , Zhigang Tu

Music stem generation, the task of producing musically-synchronized and isolated instrument audio clips, offers the potential of greater user control and better alignment with musician workflows compared to conventional text-to-music…

声音 · 计算机科学 2026-02-11 Shih-Lun Wu , Ge Zhu , Juan-Pablo Caceres , Cheng-Zhi Anna Huang , Nicholas J. Bryan

Emerging wearable devices such as smartglasses and extended reality headsets demand high-quality spatial audio capture from compact, head-worn microphone arrays. Ambisonics provides a device-agnostic spatial audio representation by mapping…

音频与语音处理 · 电气工程与系统科学 2026-01-27 Thomas Deppisch , Yang Gao , Manan Mittal , Benjamin Stahl , Christoph Hold , David Alon , Zamir Ben-Hur

The present document reviews the mathematics behind binaural rendering of sound fields that are available as spherical harmonic expansion coefficients. This process is also known as binaural ambisonic decoding. We highlight that the details…

声音 · 计算机科学 2022-09-15 Jens Ahrens

Stereophonic audio is an indispensable ingredient to enhance human auditory experience. Recent research has explored the usage of visual information as guidance to generate binaural or ambisonic audio from mono ones with stereo supervision.…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Hang Zhou , Xudong Xu , Dahua Lin , Xiaogang Wang , Ziwei Liu

Scene-based spatial audio formats, such as Ambisonics, are playback system agnostic and may therefore be favoured for delivering immersive audio experiences to a wide range of (potentially unknown) devices. The number of channels required…

音频与语音处理 · 电气工程与系统科学 2024-01-25 Christoph Hold , Leo McCormack , Archontis Politis , Ville Pulkki

Ambisonics is an established framework to capture, process, and reproduce spatial sound fields based on its spherical harmonics representation. We propose a generalization of conventional spherical ambisonics to the spheroidal coordinate…

声音 · 计算机科学 2023-01-06 Shoken Kaneko

Ambisonics Signal Matching (ASM) is a recently proposed signal-independent approach to encoding Ambisonic signal from wearable microphone arrays, enabling efficient and standardized spatial sound reproduction. However, reproduction accuracy…

音频与语音处理 · 电气工程与系统科学 2025-07-08 Yhonatan Gayer , Vladimir Tourbabin , Zamir Ben-Hur , David Alon , Boaz Rafaely

In the stereo-to-multichannel upmixing problem for music, one of the main tasks is to set the directionality of the instrument sources in the multichannel rendering results. In this paper, we propose a modified variational autoencoder model…

音频与语音处理 · 电气工程与系统科学 2022-03-24 Haici Yang , Sanna Wager , Spencer Russell , Mike Luo , Minje Kim , Wontak Kim

This paper introduces SoundSculpt, a neural network designed to extract target sound fields from ambisonic recordings. SoundSculpt employs an ambisonic-in-ambisonic-out architecture and is conditioned on both spatial information (e.g.,…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Tuochao Chen , D Shin , Hakan Erdogan , Sinan Hersek

Synthetic data generation is increasingly used in machine learning for training and data augmentation. Yet, current strategies often rely on external foundation models or datasets, whose usage is restricted in many scenarios due to policy…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Parsa Rahimi , Sebastien Marcel

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

声音 · 计算机科学 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang
‹ 上一页 1 2 3 10 下一页 ›