Related papers: Untangling Phase and Time in Monophonic Sounds
Integrated nonlinear photonic technologies, even with state-of-the-art fabrication with only a few nanometer geometry variations, face significant challenges in achieving wafer-scale yield of functional devices. A core limitation lies in…
Periodic driving of particles can create crystalline structures in their dynamics. Such systems can be used to study solid-state physics phenomena in the time domain. In addition, it is possible to realize photonic time crystals and to…
In this study we present a kernel based convolution model to characterize neural responses to natural sounds by decoding their time-varying acoustic features. The model allows to decode natural sounds from high-dimensional neural…
Neural audio codecs and autoencoders have emerged as versatile models for audio compression, transmission, feature-extraction, and latent-space generation. However, a key limitation is that most are trained to maximize reconstruction…
An imaging system is proposed for matter-wave functions that is based on producing a quadratic phase modulation on the wavefunction of a charged particle, analogous to that produced by a space or time lens. The modulation is produced by…
Synthetic dimensions in photonic structures provide unique opportunities for actively manipulating light in multiple degrees of freedom. Here, we theoretically explore a dispersive waveguide under the dynamic phase modulation that supports…
Speech separation has been very successful with deep learning techniques. Substantial effort has been reported based on approaches over spectrogram, which is well known as the standard time-and-frequency cross-domain representation for…
In kernel methods, temporal information on the data is commonly included by using time-delayed embeddings as inputs. Recently, an alternative formulation was proposed by defining a gamma-filter explicitly in a reproducing kernel Hilbert…
This paper introduces a novel data-driven strategy for synthesizing gramophone noise audio textures. A diffusion probabilistic model is applied to generate highly realistic quasiperiodic noises. The proposed model is designed to generate…
In this work, we propose a new mathematical vocoder algorithm(modified spectral inversion) that generates a waveform from acoustic features without phase estimation. The main benefit of using our proposed method is that it excludes the…
It is shown that any convolution operator in the time domain can be represented exactly as a multiplication operator in the time-scale (wavelet) domain. The Mellin transform gives a one-to-one correspondence between frequency filters…
Timbre and pitch are the two main perceptual properties of musical sounds. Depending on the target applications, we sometimes prefer to focus on one of them, while reducing the effect of the other. Researchers have managed to hand-craft…
Developing algorithms for sound classification, detection, and localization requires large amounts of flexible and realistic audio data, especially when leveraging modern machine learning and beamforming techniques. However, most existing…
Sound waves are attenuated as they propagate in amorphous materials. We investigate the mechanism driving sound attenuation in the Rayleigh scattering regime by resolving the dynamics of an excited phonon in time and space via numerical…
Consistency models have exhibited remarkable capabilities in facilitating efficient image/video generation, enabling synthesis with minimal sampling steps. It has proven to be advantageous in mitigating the computational burdens associated…
Melody is one of the most important components in music. Unlike other components in music theory, such as harmony and counterpoint, computable features for melody is urgently in need. These features are highly demanded as data-driven…
We propose a method to produce pure single photons with an arbitrary designed temporal shape in a heralded, lossless and scalable way. As the indispensable resource, the method uses pairs of time-energy entangled photons. To accomplish the…
Consistency in the output of language models is critical for their reliability and practical utility. Due to their training objective, language models learn to model the full space of possible continuations, leading to outputs that can vary…
We present a new system for simultaneous estimation of keys, chords, and bass notes from music audio. It makes use of a novel chromagram representation of audio that takes perception of loudness into account. Furthermore, it is fully based…
We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…