中文
相关论文

相关论文: Wavelet-based spatial audio framework

200 篇论文

The study of spatial audio and room acoustics aims to create immersive audio experiences by modeling the physics and psychoacoustics of how sound behaves in space. In the long history of this research area, various key technologies have…

音频与语音处理 · 电气工程与系统科学 2025-03-18 Shoichi Koyama , Enzo De Sena , Prasanga Samarasinghe , Mark R. P. Thomas , Fabio Antonacci

Neural audio codecs, neural networks which compress a waveform into discrete tokens, play a crucial role in the recent development of audio generative models. State-of-the-art codecs rely on the end-to-end training of an autoencoder and a…

声音 · 计算机科学 2025-03-26 Zineb Lahrichi , Gaëtan Hadjeres , Gael Richard , Geoffroy Peeters

We introduce Wave Arithmetic, a smooth analytical framework in which natural, integer, and rational numbers are represented not as discrete entities, but as integrals of smooth, compactly supported or periodic kernel functions. In this…

综合数学 · 数学 2025-05-27 Stanislav Semenov

We introduce phononic box crystals, namely arrays of adjoined perforated boxes, as a three-dimensional prototype for an unusual class of subwavelength metamaterials based on directly coupling resonating elements. In this case, when the…

应用物理 · 物理学 2018-08-28 Alice L. Vanel , Richard V. Craster , Ory Schnitzer

Slepian functions are orthogonal function systems that live on subdomains (for example, geographical regions on the Earth's surface, or bandlimited portions of the entire spectrum). They have been firmly established as a useful tool for the…

数值分析 · 数学 2017-11-10 Volker Michel , Frederik J. Simons

This paper introduces a new framework for supervised sound source localization referred to as virtually-supervised learning. An acoustic shoe-box room simulator is used to generate a large number of binaural single-source audio scenes.…

声音 · 计算机科学 2017-03-21 Saurabh Kataria , Clément Gaultier , Antoine Deleforge

In speech separation, time-domain approaches have successfully replaced the time-frequency domain with latent sequence feature from a learnable encoder. Conventionally, the feature is separated into speaker-specific ones at the final stage…

音频与语音处理 · 电气工程与系统科学 2026-04-01 Ui-Hyeop Shin , Sangyoun Lee , Taehan Kim , Hyung-Min Park

This paper aims to provide an unsupervised modelling approach that allows for a more flexible representation of text embeddings. It jointly encodes the words and the paragraphs as individual matrices of arbitrary column dimension with unit…

计算与语言 · 计算机科学 2022-12-01 Souvik Banerjee , Bamdev Mishra , Pratik Jawanpuria , Manish Shrivastava

The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as…

声音 · 计算机科学 2022-11-04 You Jin Kim , Hee-Soo Heo , Jee-weon Jung , Youngki Kwon , Bong-Jin Lee , Joon Son Chung

Harmonic analysis is a tool to infer cosmic topology from the measured astrophysical cosmic microwave background CMB radiation. For overall positive curvature, Platonic spherical manifolds are candidates for this analysis. We combine the…

宇宙学与河外天体物理 · 物理学 2015-06-03 Peter Kramer

Spherical microphone arrays (SMAs) and spherical loudspeaker arrays (SLAs) facilitate the study of room acoustics due to the three-dimensional analysis they provide. More recently, systems that combine both arrays, referred to as…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Hai Morgenstern , Boaz Rafaely , Markus Noisternig

Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details. Existing unified tokenizers jointly encode both in…

音频与语音处理 · 电气工程与系统科学 2026-05-28 Zhisheng Zhang , Xiang Li , Yixuan Zhou , Jing Peng , Guoyang Zeng , Zhiyong Wu

To date, various speech technology systems have adopted the vocoder approach, a method for synthesizing speech waveform that shows a major role in the performance of statistical parametric speech synthesis. WaveNet one of the best models…

声音 · 计算机科学 2021-06-15 Mohammed Salah Al-Radhi , Tamás Gábor Csapó , Csaba Zainkó , Géza Németh

Recently introduced inpainting algorithms using a combination of applied harmonic analysis and compressed sensing have turned out to be very successful. One key ingredient is a carefully chosen representation system which provides…

泛函分析 · 数学 2016-12-28 Martin Genzel , Gitta Kutyniok

3D image processing constitutes nowadays a challenging topic in many scientific fields such as medicine, computational physics and informatics. Therefore, development of suitable tools that guaranty a best treatment is a necessity.…

数值分析 · 计算机科学 2018-05-22 Malika Jallouli , Wafa Bel Hadj Khalifa , Anouar Ben Mabrouk , Mohamed Ali Mahjoub

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

声音 · 计算机科学 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

The analysis of scattering from complex objects using surface integral equations is a challenging problem. Its resolution has wide ranging applications- from crack propagation to diagnostic medicine. The two ingredients of any integral…

计算物理 · 物理学 2016-11-25 Naveen Nair , Balasubramaniam Shanker , Leo Kempel

We propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we introduce a speech-to-unit processing pipeline that captures…

音频与语音处理 · 电气工程与系统科学 2024-02-05 Jaeyeon Kim , Injune Hwang , Kyogu Lee

Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for noisy or real-world environments. Building on existing works…

计算与语言 · 计算机科学 2025-10-30 Shreyas Gopal , Ashutosh Anshul , Haoyang Li , Yue Heng Yeo , Hexin Liu , Eng Siong Chng
‹ 上一页 1 8 9 10 下一页 ›