English
Related papers

Related papers: Untangling Phase and Time in Monophonic Sounds

200 papers

Integrated nonlinear photonic technologies, even with state-of-the-art fabrication with only a few nanometer geometry variations, face significant challenges in achieving wafer-scale yield of functional devices. A core limitation lies in…

Periodic driving of particles can create crystalline structures in their dynamics. Such systems can be used to study solid-state physics phenomena in the time domain. In addition, it is possible to realize photonic time crystals and to…

In this study we present a kernel based convolution model to characterize neural responses to natural sounds by decoding their time-varying acoustic features. The model allows to decode natural sounds from high-dimensional neural…

Machine Learning · Statistics 2016-11-15 Ali Faisal , Anni Nora , Jaeho Seol , Hanna Renvall , Riitta Salmelin

Neural audio codecs and autoencoders have emerged as versatile models for audio compression, transmission, feature-extraction, and latent-space generation. However, a key limitation is that most are trained to maximize reconstruction…

Sound · Computer Science 2025-09-10 Dimitrios Bralios , Jonah Casebeer , Paris Smaragdis

An imaging system is proposed for matter-wave functions that is based on producing a quadratic phase modulation on the wavefunction of a charged particle, analogous to that produced by a space or time lens. The modulation is produced by…

Quantum Physics · Physics 2020-10-28 Brian H. Kolner

Synthetic dimensions in photonic structures provide unique opportunities for actively manipulating light in multiple degrees of freedom. Here, we theoretically explore a dispersive waveguide under the dynamic phase modulation that supports…

Optics · Physics 2025-05-16 Guangzhen Li , Danying Yu , Luqi Yuan , Xianfeng Chen

Speech separation has been very successful with deep learning techniques. Substantial effort has been reported based on approaches over spectrogram, which is well known as the standard time-and-frequency cross-domain representation for…

Sound · Computer Science 2019-04-17 Gene-Ping Yang , Chao-I Tuan , Hung-Yi Lee , Lin-shan Lee

In kernel methods, temporal information on the data is commonly included by using time-delayed embeddings as inputs. Recently, an alternative formulation was proposed by defining a gamma-filter explicitly in a reproducing kernel Hilbert…

Machine Learning · Statistics 2017-06-13 Steven Van Vaerenbergh , Simone Scardapane , Ignacio Santamaria

This paper introduces a novel data-driven strategy for synthesizing gramophone noise audio textures. A diffusion probabilistic model is applied to generate highly realistic quasiperiodic noises. The proposed model is designed to generate…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-01 Eloi Moliner , Vesa Välimäki

In this work, we propose a new mathematical vocoder algorithm(modified spectral inversion) that generates a waveform from acoustic features without phase estimation. The main benefit of using our proposed method is that it excludes the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Hyun Gon Ryu , Jeong-Hoon Kim , Simon See

It is shown that any convolution operator in the time domain can be represented exactly as a multiplication operator in the time-scale (wavelet) domain. The Mellin transform gives a one-to-one correspondence between frequency filters…

Mathematical Physics · Physics 2007-05-23 Gerald Kaiser

Timbre and pitch are the two main perceptual properties of musical sounds. Depending on the target applications, we sometimes prefer to focus on one of them, while reducing the effect of the other. Researchers have managed to hand-craft…

Sound · Computer Science 2018-11-09 Yun-Ning Hung , Yi-An Chen , Yi-Hsuan Yang

Developing algorithms for sound classification, detection, and localization requires large amounts of flexible and realistic audio data, especially when leveraging modern machine learning and beamforming techniques. However, most existing…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Luca Barbisan , Marco Levorato , Fabrizio Riente

Sound waves are attenuated as they propagate in amorphous materials. We investigate the mechanism driving sound attenuation in the Rayleigh scattering regime by resolving the dynamics of an excited phonon in time and space via numerical…

Soft Condensed Matter · Physics 2024-01-18 Shivam Mahajan , Massimo Pica Ciamarra

Consistency models have exhibited remarkable capabilities in facilitating efficient image/video generation, enabling synthesis with minimal sampling steps. It has proven to be advantageous in mitigating the computational burdens associated…

Sound · Computer Science 2024-04-23 Zhengcong Fei , Mingyuan Fan , Junshi Huang

Melody is one of the most important components in music. Unlike other components in music theory, such as harmony and counterpoint, computable features for melody is urgently in need. These features are highly demanded as data-driven…

Sound · Computer Science 2020-03-23 Zehao Wang , Shicheng Zhang , Xiaoou Chen

We propose a method to produce pure single photons with an arbitrary designed temporal shape in a heralded, lossless and scalable way. As the indispensable resource, the method uses pairs of time-energy entangled photons. To accomplish the…

Quantum Physics · Physics 2017-10-18 Valentin Averchenko , Denis Sych , Gerd Leuchs

Consistency in the output of language models is critical for their reliability and practical utility. Due to their training objective, language models learn to model the full space of possible continuations, leading to outputs that can vary…

Computation and Language · Computer Science 2025-03-04 Damien de Mijolla , Hannan Saddiq , Kim Moore

We present a new system for simultaneous estimation of keys, chords, and bass notes from music audio. It makes use of a novel chromagram representation of audio that takes perception of loudness into account. Furthermore, it is fully based…

Sound · Computer Science 2011-07-26 Yizhao Ni , Matt Mcvicar , Raul Santos-Rodriguez , Tijl De Bie

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak