Related papers: Untangling Phase and Time in Monophonic Sounds
The advent of novel nonlinear materials has stirred unprecedented interest in exploring the use of temporal inhomogeneities to achieve novel forms of wave control, amidst the greater vision of engineering metamaterials across both space and…
Synthetic magnetism has been recently realized using spatiotemporal modulation patterns, producing non-reciprocal steering of charge-neutral particles such as photons and phonons. Here, we design and experimentally demonstrate a…
We present a novel model designed for resource-efficient multichannel speech enhancement in the time domain, with a focus on low latency, lightweight, and low computational requirements. The proposed model incorporates explicit spatial and…
Acoustic holograms have promising applications in sound-field reconstruction, particle manipulation, ultrasonic haptics and therapy. This paper reports on the theoretical, numerical, and experimental investigation of multiplexed acoustic…
FM Synthesis is a well-known algorithm used to generate complex timbre from a compact set of design primitives. Typically featuring a MIDI interface, it is usually impractical to control it from an audio source. On the other hand,…
Acoustic metasurfaces manipulate waves with specially designed structures and achieve properties that natural materials cannot offer. Similar surfaces work in audio frequency range as well and lead to marvelous acoustic phenomena that can…
Mandarin Chinese is characterized by being a tonal language; the pitch (or $F_0$) of its utterances carries considerable linguistic information. However, speech samples from different individuals are subject to changes in amplitude and…
Convolutional layers with 1-D filters are often used as frontend to encode audio signals. Unlike fixed time-frequency representations, they can adapt to the local characteristics of input data. However, 1-D filters on raw audio are hard to…
Video to sound generation aims to generate realistic and natural sound given a video input. However, previous video-to-sound generation methods can only generate a random or average timbre without any controls or specializations of the…
Quantum manipulation of individual phonons could offer new resources for studying fundamental physics and creating an innovative platform in quantum information science. Here, we propose to generate quantum states of strongly correlated…
Phase retrieval is in general a non-convex and non-linear task and the corresponding algorithms struggle with the issue of local minima. We consider the case where the measurement samples within typically very small and disconnected subsets…
We introduce VampNet, a masked acoustic token modeling approach to music synthesis, compression, inpainting, and variation. We use a variable masking schedule during training which allows us to sample coherent music from the model by…
This paper introduces an extendable modular system that compiles a range of music feature extraction models to aid music information retrieval research. The features include musical elements like key, downbeats, and genre, as well as audio…
The unique conduction properties of condensed matter systems with topological order have recently inspired a quest for similar effects in classical wave phenomena. Acoustic topological insulators, in particular, hold the promise to…
We consider the resonance and scattering properties of a composite medium containing scatterers whose properties are modulated in time. When excited with an incident wave of a single frequency, the scattered field consists of a family of…
Speech enhancement is crucial for ubiquitous human-computer interaction. Recently, ultrasound-based acoustic sensing has emerged as an attractive choice for speech enhancement because of its superior ubiquity and performance. However, due…
Historically, most speech models in machine-learning have used the mel-spectrogram as a speech representation. Recently, discrete audio tokens produced by neural audio codecs have become a popular alternate speech representation for speech…
In this paper, a deep-learning-based method for sound field reconstruction is proposed. It is shown the possibility to reconstruct the magnitude of the sound pressure in the frequency band 30-300 Hz for an entire room by using a very low…
This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture of two sources, with their…
Voice Conversion (VC) converts the voice of a source speech to that of a target while maintaining the source's content. Speech can be mainly decomposed into four components: content, timbre, rhythm and pitch. Unfortunately, most related…