Related papers: Untangling Phase and Time in Monophonic Sounds
Decomposition of an audio mixture into harmonic and percussive components, namely harmonic/percussive source separation (HPSS), is a useful pre-processing tool for many audio applications. Popular approaches to HPSS exploit the distinctive…
Sampling in shift-invariant spaces is a realistic model for signals with smooth spectrum. In this paper, we consider phaseless sampling and reconstruction of real-valued signals in a shift-invariant space from their magnitude measurements…
We tackle the task of conditional music generation. We introduce MusicGen, a single Language Model (LM) that operates over several streams of compressed discrete music representation, i.e., tokens. Unlike prior work, MusicGen is comprised…
The reconstruction mechanisms built by the human auditory system during sound reconstruction are still a matter of debate. The purpose of this study is to refine the auditory cortex model introduced in [9], and inspired by the geometrical…
This article explains how to apply time--frequency scattering, a convolutional operator extracting modulations in the time--frequency domain at different rates and scales, to the re-synthesis and manipulation of audio textures. After…
We propose a simple phenomenological model describing composite crystals, constructed from two parallel sets of periodic inter-penetrating chains. In the harmonic approximation and neglecting thermal fluctuations we find the eigenmodes of…
Systems for synthesizer sound matching, which automatically set the parameters of a synthesizer to emulate an input sound, have the potential to make the process of synthesizer programming faster and easier for novice and experienced…
We explore a novel way of conceptualising the task of polyphonic music transcription, using so-called invertible neural networks. Invertible models unify both discriminative and generative aspects in one function, sharing one set of…
Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in the perception of speaker identity. To bridge this gap, we…
Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthesis models currently rely on annotated audio data, but it is…
We demonstrate how conditional generation from diffusion models can be used to tackle a variety of realistic tasks in the production of music in 44.1kHz stereo audio with sampling-time guidance. The scenarios we consider include…
Sampling is classically performed by recording the amplitude of an input signal at given time instants; however, sampling and reconstructing a signal using multiple devices in parallel becomes a more difficult problem to solve when the…
The task of bandwidth extension addresses the generation of missing high frequencies of audio signals based on knowledge of the low-frequency part of the sound. This task applies to various problems, such as audio coding or audio…
Spatial sampling is traditionally studied in a static setting where static sensors scattered around space take measurements of the spatial field at their locations. In this paper we study the emerging paradigm of sampling and reconstructing…
This paper focuses on phase retrieval from phaseless total-field data in biharmonic scattering problems. We prove that a phased biharmonic wave can be uniquely determined by the modulus of the total biharmonic wave within a nonempty domain.…
Continuous-variable encoding of quantum information in the optical domain has recently yielded large temporal and spectral entangled states instrumental for quantum computing and quantum communication. We introduce a protocol for the…
Reverberant sound fields are often modeled as isotropic. However, it has been observed that spatial properties change during the decay of the sound field energy, due to non-isotropic attenuation in non-ideal rooms. In this letter, a model…
A music mashup combines audio elements from two or more songs to create a new work. To reduce the time and effort required to make them, researchers have developed algorithms that predict the compatibility of audio elements. Prior work has…
In many dynamical probes of a quantum system, quite often multiple eigenmodes are excited. Therefore, the experimental data can be quite messy due to the mixing of different modes, as well as the background noise, despite that each mode…
Metamaterials can enable peculiar static and dynamic behavior (such as negative effective mass density, dynamical stiffness, and Poisson's ratio) due to their geometry rather than their chemical composition. The geometry of these…