English
Related papers

Related papers: Wavelet-based spatial audio framework

200 papers

The study of spatial audio and room acoustics aims to create immersive audio experiences by modeling the physics and psychoacoustics of how sound behaves in space. In the long history of this research area, various key technologies have…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-18 Shoichi Koyama , Enzo De Sena , Prasanga Samarasinghe , Mark R. P. Thomas , Fabio Antonacci

Neural audio codecs, neural networks which compress a waveform into discrete tokens, play a crucial role in the recent development of audio generative models. State-of-the-art codecs rely on the end-to-end training of an autoencoder and a…

Sound · Computer Science 2025-03-26 Zineb Lahrichi , Gaëtan Hadjeres , Gael Richard , Geoffroy Peeters

We introduce Wave Arithmetic, a smooth analytical framework in which natural, integer, and rational numbers are represented not as discrete entities, but as integrals of smooth, compactly supported or periodic kernel functions. In this…

General Mathematics · Mathematics 2025-05-27 Stanislav Semenov

We introduce phononic box crystals, namely arrays of adjoined perforated boxes, as a three-dimensional prototype for an unusual class of subwavelength metamaterials based on directly coupling resonating elements. In this case, when the…

Applied Physics · Physics 2018-08-28 Alice L. Vanel , Richard V. Craster , Ory Schnitzer

Slepian functions are orthogonal function systems that live on subdomains (for example, geographical regions on the Earth's surface, or bandlimited portions of the entire spectrum). They have been firmly established as a useful tool for the…

Numerical Analysis · Mathematics 2017-11-10 Volker Michel , Frederik J. Simons

This paper introduces a new framework for supervised sound source localization referred to as virtually-supervised learning. An acoustic shoe-box room simulator is used to generate a large number of binaural single-source audio scenes.…

Sound · Computer Science 2017-03-21 Saurabh Kataria , Clément Gaultier , Antoine Deleforge

In speech separation, time-domain approaches have successfully replaced the time-frequency domain with latent sequence feature from a learnable encoder. Conventionally, the feature is separated into speaker-specific ones at the final stage…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-01 Ui-Hyeop Shin , Sangyoun Lee , Taehan Kim , Hyung-Min Park

This paper aims to provide an unsupervised modelling approach that allows for a more flexible representation of text embeddings. It jointly encodes the words and the paragraphs as individual matrices of arbitrary column dimension with unit…

Computation and Language · Computer Science 2022-12-01 Souvik Banerjee , Bamdev Mishra , Pratik Jawanpuria , Manish Shrivastava

The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as…

Sound · Computer Science 2022-11-04 You Jin Kim , Hee-Soo Heo , Jee-weon Jung , Youngki Kwon , Bong-Jin Lee , Joon Son Chung

Harmonic analysis is a tool to infer cosmic topology from the measured astrophysical cosmic microwave background CMB radiation. For overall positive curvature, Platonic spherical manifolds are candidates for this analysis. We combine the…

Cosmology and Nongalactic Astrophysics · Physics 2015-06-03 Peter Kramer

Spherical microphone arrays (SMAs) and spherical loudspeaker arrays (SLAs) facilitate the study of room acoustics due to the three-dimensional analysis they provide. More recently, systems that combine both arrays, referred to as…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Hai Morgenstern , Boaz Rafaely , Markus Noisternig

Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details. Existing unified tokenizers jointly encode both in…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-28 Zhisheng Zhang , Xiang Li , Yixuan Zhou , Jing Peng , Guoyang Zeng , Zhiyong Wu

To date, various speech technology systems have adopted the vocoder approach, a method for synthesizing speech waveform that shows a major role in the performance of statistical parametric speech synthesis. WaveNet one of the best models…

Sound · Computer Science 2021-06-15 Mohammed Salah Al-Radhi , Tamás Gábor Csapó , Csaba Zainkó , Géza Németh

Recently introduced inpainting algorithms using a combination of applied harmonic analysis and compressed sensing have turned out to be very successful. One key ingredient is a carefully chosen representation system which provides…

Functional Analysis · Mathematics 2016-12-28 Martin Genzel , Gitta Kutyniok

3D image processing constitutes nowadays a challenging topic in many scientific fields such as medicine, computational physics and informatics. Therefore, development of suitable tools that guaranty a best treatment is a necessity.…

Numerical Analysis · Computer Science 2018-05-22 Malika Jallouli , Wafa Bel Hadj Khalifa , Anouar Ben Mabrouk , Mohamed Ali Mahjoub

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

Sound · Computer Science 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

The analysis of scattering from complex objects using surface integral equations is a challenging problem. Its resolution has wide ranging applications- from crack propagation to diagnostic medicine. The two ingredients of any integral…

Computational Physics · Physics 2016-11-25 Naveen Nair , Balasubramaniam Shanker , Leo Kempel

We propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we introduce a speech-to-unit processing pipeline that captures…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-05 Jaeyeon Kim , Injune Hwang , Kyogu Lee

Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for noisy or real-world environments. Building on existing works…

Computation and Language · Computer Science 2025-10-30 Shreyas Gopal , Ashutosh Anshul , Haoyang Li , Yue Heng Yeo , Hexin Liu , Eng Siong Chng
‹ Prev 1 8 9 10 Next ›