English
Related papers

Related papers: SpecTNT: a Time-Frequency Transformer for Music Au…

200 papers

Transformers and State-Space Models (SSMs) have advanced audio classification by modeling spectrograms as sequences of patches. However, existing models such as the Audio Spectrogram Transformer (AST) and Audio Mamba (AuM) adopt square…

Sound · Computer Science 2025-09-01 Aditya Makineni , Baocheng Geng , Qing Tian

Audio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class information of audio events, and the order in which they occur…

Sound · Computer Science 2022-10-25 Yuanbo Hou , Yun Wang , Wenwu Wang , Dick Botteldooren

Intelligent spectrum management is crucial for improving spectrum efficiency and achieving secure utilization of spectrum resources. However, existing intelligent spectrum management methods, typically based on small-scale models, suffer…

Signal Processing · Electrical Eng. & Systems 2025-12-16 Fuhui Zhou , Chunyu Liu , Hao Zhang , Wei Wu , Qihui Wu , Tony Q. S. Quek , Chan-Byoung Chae

Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual signals. The core challenge is to effectively combine temporal numerical patterns with the context…

Machine Learning · Computer Science 2026-02-04 Huu Hiep Nguyen , Minh Hoang Nguyen , Dung Nguyen , Hung Le

In time series classification and regression, signals are typically mapped into some intermediate representation used for constructing models. Since the underlying task is often insensitive to time shifts, these representations are required…

Sound · Computer Science 2019-07-16 Joakim Andén , Vincent Lostanlen , Stéphane Mallat

We present a framework based on neural networks to extract music scores directly from polyphonic audio in an end-to-end fashion. Most previous Automatic Music Transcription (AMT) methods seek a piano-roll representation of the pitches, that…

Sound · Computer Science 2019-10-29 Miguel A. Román , Antonio Pertusa , Jorge Calvo-Zaragoza

Robustness against temporal variations is important for emotion recognition from speech audio, since emotion is ex-pressed through complex spectral patterns that can exhibit significant local dilation and compression on the time axis…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-10 Eric Guizzo , Tillman Weyde , Jack Barnett Leveson

We introduce the joint time-frequency scattering transform, a time shift invariant descriptor of time-frequency structure for audio classification. It is obtained by applying a two-dimensional wavelet transform in time and log-frequency to…

Sound · Computer Science 2018-08-06 Joakim Andén , Vincent Lostanlen , Stéphane Mallat

Automatic music transcription (AMT) aims to convert raw audio to symbolic music representation. As a fundamental problem of music information retrieval (MIR), AMT is considered a difficult task even for trained human experts due to overlap…

Sound · Computer Science 2023-02-28 Shenli Yuan , Lingjie Kong , Jiushuang Guo

Audio captioning aims to automatically generate a natural language description of an audio clip. Most captioning models follow an encoder-decoder architecture, where the decoder predicts words based on the audio features extracted by the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-22 Xinhao Mei , Xubo Liu , Qiushi Huang , Mark D. Plumbley , Wenwu Wang

We present a transformer-based speech-declipping model that effectively recovers clipped signals across a wide range of input signal-to-distortion ratios (SDRs). While recent time-domain deep neural network (DNN)-based declippers have…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Younghoo Kwon , Jung-Woo Choi

Most recent research about automatic music transcription (AMT) uses convolutional neural networks and recurrent neural networks to model the mapping from music signals to symbolic notation. Based on a high-resolution piano transcription…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-11 Longshen Ou , Ziyi Guo , Emmanouil Benetos , Jiqing Han , Ye Wang

Automatic Music Transcription (AMT), inferring musical notes from raw audio, is a challenging task at the core of music understanding. Unlike Automatic Speech Recognition (ASR), which typically focuses on the words of a single speaker, AMT…

Sound · Computer Science 2022-03-16 Josh Gardner , Ian Simon , Ethan Manilow , Curtis Hawthorne , Jesse Engel

An effective understanding of the contextual environment and accurate motion forecasting of surrounding agents is crucial for the development of autonomous vehicles and social mobile robots. This task is challenging since the behavior of an…

Computer Vision and Pattern Recognition · Computer Science 2021-06-08 Defu Cao , Jiachen Li , Hengbo Ma , Masayoshi Tomizuka

We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that require large model capacity and substantial memory…

Sound · Computer Science 2025-03-24 Tao Feng , Zhiyuan Zhao , Yifan Xie , Yuqi Ye , Xiangyang Luo , Xun Guan , Yu Li

Current and upcoming generations of visible-shortwave infrared (VSWIR) imaging spectrometers promise unprecedented capacity to quantify Earth System processes across the globe. However, reliable cloud screening remains a fundamental…

Machine Learning · Computer Science 2025-07-08 Jake H. Lee , Michael Kiper , David R. Thompson , Philip G. Brodrick

Following the successful application of vision transformers in multiple computer vision tasks, these models have drawn the attention of the signal processing community. This is because signals are often represented as spectrograms (e.g.…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Nicolae-Catalin Ristea , Radu Tudor Ionescu , Fahad Shahbaz Khan

We introduce the Latent Fourier Transform (LatentFT), a framework that provides novel frequency-domain controls for generative music models. LatentFT combines a diffusion autoencoder with a latent-space Fourier transform to separate musical…

Sound · Computer Science 2026-04-21 Mason Wang , Cheng-Zhi Anna Huang

This paper introduces a novel Token-and-Duration Transducer (TDT) architecture for sequence-to-sequence tasks. TDT extends conventional RNN-Transducer architectures by jointly predicting both a token and its duration, i.e. the number of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-31 Hainan Xu , Fei Jia , Somshubra Majumdar , He Huang , Shinji Watanabe , Boris Ginsburg

Targeting at both high efficiency and performance, we propose AlignTTS to predict the mel-spectrum in parallel. AlignTTS is based on a Feed-Forward Transformer which generates mel-spectrum from a sequence of characters, and the duration of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-05 Zhen Zeng , Jianzong Wang , Ning Cheng , Tian Xia , Jing Xiao