English
Related papers

Related papers: Time-Frequency Phase Retrieval for Audio -- The Ef…

200 papers

Spectral interference, the frequency counterpart of the beating phenomenon in the time domain, can severely distort time-frequency representations (TFRs) in physical applications. We study this phenomenon for the short-time Fourier…

Classical Analysis and ODEs · Mathematics 2026-01-19 Shrikant Chand , James Nolen , Hau-Tieng Wu

Speech super-resolution (SSR) enhances low-resolution speech by increasing the sampling rate. While most SSR methods focus on magnitude reconstruction, recent research highlights the importance of phase reconstruction for improved…

Convolutional neural networks (CNN) are widely used for speech emotion recognition (SER). In such cases, the short time fourier transform (STFT) spectrogram is the most popular choice for representing speech, which is fed as input to the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-09 Shruti Gupta , Md. Shah Fahad , Akshay Deepak

Multi-frame algorithms for single-microphone speech enhancement, e.g., the multi-frame minimum variance distortionless response (MFMVDR) filter, are able to exploit speech correlation across adjacent time frames in the short-time Fourier…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-17 Marvin Tammen , Simon Doclo

The reconstruction of a function from its spectrogram (i.e., the absolute value of its short-time Fourier transform (STFT)) arises as a key problem in several important applications, including coherent diffraction imaging and audio…

Functional Analysis · Mathematics 2023-10-02 Philipp Grohs , Lukas Liehr

We investigate the uniqueness of short-time Fourier transform phase retrieval problems in $L^2(\mathbb{R})$. In particular, for underlying window functions whose Fourier transform decay faster than any exponential function, we derive…

Functional Analysis · Mathematics 2025-11-21 Shuang Guan , Kasso A. Okoudjou

Parameter-efficient transfer learning (PETL) methods have emerged as a solid alternative to the standard full fine-tuning approach. They only train a few extra parameters for each downstream task, without sacrificing performance and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-16 Umberto Cappellazzo , Daniele Falavigna , Alessio Brutti , Mirco Ravanelli

From the existing research it has been observed that many techniques and methodologies are available for performing every step of Automatic Speech Recognition (ASR) system, but the performance (Minimization of Word Error Recognition-WER and…

Computation and Language · Computer Science 2013-03-25 Urmila Shrawankar , Vilas Thakare

This paper explores the innovative application of the Fractional Fourier Transform (FrFT) in sound synthesis, highlighting its potential to redefine time-frequency analysis in audio processing. As an extension of the classical Fourier…

Sound · Computer Science 2025-06-12 Esteban Gutiérrez , Rodrigo Cádiz , Carlos Sing Long , Frederic Font , Xavier Serra

We describe a new algorithm to solve a particular phase retrieval problem, that has wide applications in audio processing: the reconstruction of a function from its scalogram, that is from the modulus of its wavelet transform. It is a…

Optimization and Control · Mathematics 2017-04-11 Irène Waldspurger

Recent advancements in video restoration have focused on recovering high-quality video frames from low-quality inputs. Compared with static images, the performance of video restoration significantly depends on efficient exploitation of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Ranjith Merugu , Mohammad Sameer Suhail , Akshay P Sarashetti , Venkata Bharath Reddy Reddem , Pankaj Kumar Bajpai , Amit Satish Unde

This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-31 Shinji Takaki , Toru Nakashika , Xin Wang , Junichi Yamagishi

In this paper, we revisit the use of spectrograms in neural networks, by making the window length a continuous parameter optimizable by gradient descent instead of an empirically tuned integer-valued hyperparameter. The contribution is…

Machine Learning · Computer Science 2022-08-26 Maxime Leiber , Axel Barrau , Yosra Marnissi , Dany Abboud

Deep learning-based techniques for automatic dysarthric speech detection have recently attracted interest in the research community. State-of-the-art techniques typically learn neurotypical and dysarthric discriminative representations by…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-04 Ina Kodrasi

Time Series Foundation Models (TSFMs) have borrowed the long context paradigm from natural language processing under the premise that feeding more history into the model improves forecast quality. But in stochastic domains, distant history…

Machine Learning · Computer Science 2026-05-12 Rishi Ahuja , Kumar Prateek , Simranjit Singh , Vijay Kumar

Recent studies applied Parameter Efficient Fine-Tuning techniques (PEFTs) to efficiently narrow the performance gap between pre-training and downstream. There are two important factors for various PEFTs, namely, the accessible data size and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Yuxin Tian , Mouxing Yang , Yunfan Li , Dayiheng Liu , Xingzhang Ren , Xi Peng , Jiancheng Lv

Audio Super-Resolution (SR) is an important topic as low-resolution recordings are ubiquitous in daily life. In this paper, we focus on the music SR task, which is challenging due to the wide frequency response and dynamic range of music.…

Sound · Computer Science 2024-02-20 Yenan Zhang , Guilly Kolkman , Hiroshi Watanabe

Diffusion models have demonstrated remarkable success in generative tasks, including audio super-resolution (SR). In many applications like movie post-production and album mastering, substantial computational budgets are available for…

Sound · Computer Science 2025-08-05 Yizhu Jin , Zhen Ye , Zeyue Tian , Haohe Liu , Qiuqiang Kong , Yike Guo , Wei Xue

Dynamic parameterization of acoustic environments has drawn widespread attention in the field of audio processing. Precise representation of local room acoustic characteristics is crucial when designing audio filters for various audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-26 Chunxi Wang , Maoshen Jia , Meiran Li , Changchun Bao , Wenyu Jin

Radio frequency fingerprint (RFF) identification technology, which exploits relatively stable hardware imperfections, is highly susceptible to constantly changing channel effects. Although various channel-robust RFF feature extraction…

Signal Processing · Electrical Eng. & Systems 2026-02-10 Xuan Yang , Dongming Li , Yi Lou , Xianglin Fan