English
Related papers

Related papers: Joint Time-Frequency Scattering for Audio Classifi…

200 papers

A finite-energy signal is represented by a square-integrable, complex-valued function $t\mapsto s(t)$ of a real variable $t$, interpreted as time. Similarly, a noisy signal is represented by a random process. Time-frequency analysis, a…

Signal Processing · Electrical Eng. & Systems 2025-04-16 Barbara Pascal , Rémi Bardenet

Sequential audio event tagging can provide not only the type information of audio events, but also the order information between events and the number of events that occur in an audio clip. Most previous works on audio event sequence…

Sound · Computer Science 2022-03-23 Yuanbo Hou , Zhaoyi Liu , Bo Kang , Yun Wang , Dick Botteldooren

Recently it has been shown that the intensity time-bandwidth product of optical signals can be engineered to match that of the data acquisition instrument. In particular, it is possible to slow down an ultrafast signal, resulting in…

Optics · Physics 2015-06-22 Jacky Chan , Ata Mahjoubfar , Mohammad H. Asghari , Bahram Jalali

The paper focuses on inpainting missing parts of an audio signal spectrogram, i.e., estimating the lacking time-frequency coefficients. The autoregression-based Janssen algorithm, a state-of-the-art for the time-domain audio inpainting, is…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-09 Ondřej Mokrý , Peter Balušík , Pavel Rajmic

Time-series classification is an important domain of machine learning and a plethora of methods have been developed for the task. In comparison to existing approaches, this study presents a novel method which decomposes a time-series…

Machine Learning · Computer Science 2015-03-12 Josif Grabocka , Lars Schmidt-Thieme

We study detection and imaging of small reflectors in heavy clutter, using an array of transducers that emits and receives sound waves. Heavy clutter means that multiple scattering of the waves in the heterogeneous host medium is strong and…

Computational Physics · Physics 2018-06-27 Liliana Borcea , George Papanicolaou , Chrysoula Tsogka

With the emergence of GAN-based vocoders, the discriminator, as a crucial component, has been developed recently. In our work, we focus on improving the time-frequency based discriminator. Particularly, Short-Time Fourier Transform (STFT)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-04 Nan Xu , Zhaolong Huang , Xiao Zeng

A novel approach to improving the performances of confocal scanning imaging is proposed. We experimentally demonstrate its feasibility using acoustic waves. It relies on a new way to encode spatial information using the temporal dimension.…

This letter introduces an innovative method to enhance the quality of audio time stretching by precisely decomposing a sound into sines, transients, and noise and by improving the processing of the latter component. While there are…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-25 Eloi Moliner , Leonardo Fierro , Alec Wright , Matti Hämäläinen , Vesa Välimäki

Assessment of voice signals has long been performed with the assumption of periodicity as this facilitates analysis. Near periodicity of normal voice signals makes short-time harmonic modeling an appealing choice to extract vocal feature…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-10 Takeshi Ikuma , Andrew J. McWhorter , Lacey Adkins , Melda Kunduk

The finite STFT Synchrosqueezing transform is a time-frequency analysis method that can decompose finite complex signals into time-varying oscillatory components. This representation is sparse and invertible, allowing recovery of the…

Numerical Analysis · Mathematics 2017-09-26 Mozhgan Mohammadpour , Bastiaan Kleijn , Rajab Ali Kamyabi Gol

In deep time series forecasting, the Fourier Transform (FT) is extensively employed for frequency representation learning. However, it often struggles in capturing multi-scale, time-sensitive patterns. Although the Wavelet Transform (WT)…

Machine Learning · Computer Science 2026-02-09 Ziyu Zhou , Jiaxi Hu , Qingsong Wen , James T. Kwok , Yuxuan Liang

A fundamental problem in wireless communication is the time-frequency shift (TFS) problem: Find the time-frequency shift of a signal in a noisy environment. The shift is the result of time asynchronization of a sender with a receiver, and…

Information Theory · Computer Science 2011-12-22 Alexander Fish , Shamgar Gurevich , Ronny Hadani , Akbar Sayeed , Oded Schwartz

In this article we study the inverse problem of thermoacoustic tomography (TAT) on a medium with attenuation represented by a time- convolution (or memory) term, and whose consideration is motivated by the modeling of ultrasound waves in…

Analysis of PDEs · Mathematics 2017-03-29 Sebastian Acosta , Benjamin Palacios

In this work, we focus on front-end design for speech deepfake detectors, the component that determines the discriminative acoustic cues provided to the classifier. Existing approaches are primarily categorized into two types. Hand-crafted…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-01 Xi Xuan , Davide Carbone , Wenxin Zhang , Ruchi Pandey , Tomi H. Kinnunen

Extracting features from the speech is the most critical process in speech signal processing. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used features in the majority of the speaker and speech recognition applications,…

Sound · Computer Science 2025-10-31 Rinku Sebastian , Simon O'Keefe , Martin Trefzer

A complete framework for the linear time-invariant (LTI) filtering theory of bivariate signals is proposed based on a tailored quaternion Fourier transform. This framework features a direct description of LTI filters in terms of their…

Signal Processing · Electrical Eng. & Systems 2018-08-29 Julien Flamant , Pierre Chainais , Nicolas Le Bihan

The short-time Fourier transform (STFT) usually computes the same number of frequency components as the frame length while overlapping adjacent time frames by more than half. As a result, the number of components of a spectrogram matrix…

Signal Processing · Electrical Eng. & Systems 2020-10-29 Daichi Kitahara

Commonly used features in spoken language identification (LID), such as mel-spectrogram or MFCC, lose high-frequency information due to windowing. The loss further increases for longer temporal contexts. To improve generalization of the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-04 Spandan Dey , Premjeet Singh , Goutam Saha

In the process of recording, storage and transmission of time-domain audio signals, errors may be introduced that are difficult to correct in an unsupervised way. Here, we train a convolutional deep neural network to re-synthesize input…

Sound · Computer Science 2015-03-20 Andrew J. R. Simpson