English
Related papers

Related papers: Untangling Phase and Time in Monophonic Sounds

200 papers

We introduce a new audio processing technique that increases the sampling rate of signals such as speech or music using deep convolutional neural networks. Our model is trained on pairs of low and high-quality audio examples; at test-time,…

Sound · Computer Science 2017-08-03 Volodymyr Kuleshov , S. Zayd Enam , Stefano Ermon

Audio inpainting, i.e., the task of restoring missing or occluded audio signal samples, usually relies on sparse representations or autoregressive modeling. In this paper, we propose to structure the spectrogram with nonnegative matrix…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-06 Ondřej Mokrý , Paul Magron , Thomas Oberlin , Cédric Févotte

We propose a method for filling gaps and removing interferences in time series for applications involving continuous monitoring of environmental variables. The approach is non-parametric and based on an iterative pattern-matching between…

Geophysics · Physics 2015-08-11 Gregoire Mariethoz , Niklas Linde , Damien Jougnot , Hassan Rezaee

Models for audio source separation usually operate on the magnitude spectrum, which ignores phase information and makes separation performance dependant on hyper-parameters for the spectral front-end. Therefore, we investigate end-to-end…

Sound · Computer Science 2018-06-11 Daniel Stoller , Sebastian Ewert , Simon Dixon

High-quality audio is essential in a wide range of applications, including online communication, virtual assistants, and the multimedia industry. However, degradation caused by noise, compression, and transmission artifacts remains a major…

The effective medium representation is fundamental in providing a performance-to-design approach for many devices based on metamaterials. While there are recent works in extending the effective medium concept into the temporal domain, a…

Applied Physics · Physics 2021-08-25 Xinhua Wen , Xinghong Zhu , Hong Wei Wu , Jensen Li

Speech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm. Obtaining disentangled representations of these components is useful in many speech analysis and generation applications.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-16 Kaizhi Qian , Yang Zhang , Shiyu Chang , David Cox , Mark Hasegawa-Johnson

Over the last couple of years, the digital coding acoustic metasurfaces have been developed rapidly as a highly active research area for their unique and flexible manipulation of acoustic wavefronts. Nevertheless, all recent attentions in…

Applied Physics · Physics 2020-03-31 Hamid Rajabalipanah , Mohammad Hosein Fakheri , Ali Abdolali

In deep learning research, many melody extraction models rely on redesigning neural network architectures to improve performance. In this paper, we propose an input feature modification and a training objective modification based on two…

Sound · Computer Science 2023-08-08 Keren Shao , Ke Chen , Taylor Berg-Kirkpatrick , Shlomo Dubnov

We add a time-dependent potential to the inhomogeneous wave equation and consider the task of reconstructing this potential from measurements of the wave field. This dynamic inverse problem becomes more involved compared to static…

Numerical Analysis · Mathematics 2017-08-24 Thies Gerken , Armin Lechleiter

Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. Current methods either rely on disjoint pipelines, which fail to capture fine-grained…

Sound · Computer Science 2026-05-28 Yuyue Wang , Xihua Wang , Xin Cheng , Yijing Chen , Ruihua Song

An acoustic reverberator consisting of a network of delay lines connected via scattering junctions is proposed. All parameters of the reverberator are derived from physical properties of the enclosure it simulates. It allows for simulation…

Sound · Computer Science 2015-07-10 Enzo De Sena , Huseyin Hacihabiboglu , Zoran Cvetkovic , Julius O. Smith

A key aspect of machine learning models lies in their ability to learn efficient intermediate features. However, the input representation plays a crucial role in this process, and polyphonic musical scores remain a particularly complex type…

Machine Learning · Computer Science 2021-09-09 Mathieu Prang , Philippe Esling

The recent success of raw audio waveform synthesis models like WaveNet motivates a new approach for music synthesis, in which the entire process --- creating audio samples from a score and instrument information --- is modeled using…

Sound · Computer Science 2018-11-02 Jong Wook Kim , Rachel Bittner , Aparna Kumar , Juan Pablo Bello

Nonnegative Matrix Factorization (NMF) is a powerful tool for decomposing mixtures of audio signals in the Time-Frequency (TF) domain. In applications such as source separation, the phase recovery for each extracted component is a major…

Sound · Computer Science 2016-11-17 Paul Magron , Roland Badeau , Bertrand David

Broadband temporal modes of pulsed optical fields have been recently recognized as very promising for photonic quantum information processing and time-frequency metrology. Exploiting their full potential demands efficient and flexible tools…

Optics · Physics 2025-08-08 B Dioum , S Srivastava , M Karpiński , G Patera

In time series classification and regression, signals are typically mapped into some intermediate representation used for constructing models. Since the underlying task is often insensitive to time shifts, these representations are required…

Sound · Computer Science 2019-07-16 Joakim Andén , Vincent Lostanlen , Stéphane Mallat

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

Sound · Computer Science 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin

The intrinsic limitation of the material nonlinearity inevitably results in the poor mode purity, conversion efficiency and real-time reconfigurability of the generated harmonic waves, both in optics and acoustics. Rotational Doppler effect…

Applied Physics · Physics 2022-05-02 Chengbo Hu , Wei Wang , Jincheng Ni , Yujiang Ding , Jingkai Weng , Bin Liang , Cheng-Wei Qiu , Jianchun Cheng

Single-channel speech separation in time domain and frequency domain has been widely studied for voice-driven applications over the past few years. Most of previous works assume known number of speakers in advance, however, which is not…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-02 Yiming Xiao , Haijian Zhang