Related papers: Wavelet-based spatial audio framework
The study of spatial audio and room acoustics aims to create immersive audio experiences by modeling the physics and psychoacoustics of how sound behaves in space. In the long history of this research area, various key technologies have…
Neural audio codecs, neural networks which compress a waveform into discrete tokens, play a crucial role in the recent development of audio generative models. State-of-the-art codecs rely on the end-to-end training of an autoencoder and a…
We introduce Wave Arithmetic, a smooth analytical framework in which natural, integer, and rational numbers are represented not as discrete entities, but as integrals of smooth, compactly supported or periodic kernel functions. In this…
We introduce phononic box crystals, namely arrays of adjoined perforated boxes, as a three-dimensional prototype for an unusual class of subwavelength metamaterials based on directly coupling resonating elements. In this case, when the…
Slepian functions are orthogonal function systems that live on subdomains (for example, geographical regions on the Earth's surface, or bandlimited portions of the entire spectrum). They have been firmly established as a useful tool for the…
This paper introduces a new framework for supervised sound source localization referred to as virtually-supervised learning. An acoustic shoe-box room simulator is used to generate a large number of binaural single-source audio scenes.…
In speech separation, time-domain approaches have successfully replaced the time-frequency domain with latent sequence feature from a learnable encoder. Conventionally, the feature is separated into speaker-specific ones at the final stage…
This paper aims to provide an unsupervised modelling approach that allows for a more flexible representation of text embeddings. It jointly encodes the words and the paragraphs as individual matrices of arbitrary column dimension with unit…
The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as…
Harmonic analysis is a tool to infer cosmic topology from the measured astrophysical cosmic microwave background CMB radiation. For overall positive curvature, Platonic spherical manifolds are candidates for this analysis. We combine the…
Spherical microphone arrays (SMAs) and spherical loudspeaker arrays (SLAs) facilitate the study of room acoustics due to the three-dimensional analysis they provide. More recently, systems that combine both arrays, referred to as…
Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details. Existing unified tokenizers jointly encode both in…
To date, various speech technology systems have adopted the vocoder approach, a method for synthesizing speech waveform that shows a major role in the performance of statistical parametric speech synthesis. WaveNet one of the best models…
Recently introduced inpainting algorithms using a combination of applied harmonic analysis and compressed sensing have turned out to be very successful. One key ingredient is a carefully chosen representation system which provides…
3D image processing constitutes nowadays a challenging topic in many scientific fields such as medicine, computational physics and informatics. Therefore, development of suitable tools that guaranty a best treatment is a necessity.…
Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…
How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…
The analysis of scattering from complex objects using surface integral equations is a challenging problem. Its resolution has wide ranging applications- from crack propagation to diagnostic medicine. The two ingredients of any integral…
We propose a framework to learn semantics from raw audio signals using two types of representations, encoding contextual and phonetic information respectively. Specifically, we introduce a speech-to-unit processing pipeline that captures…
Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for noisy or real-world environments. Building on existing works…