English
Related papers

Related papers: Multitrack Music Transcription with a Time-Frequen…

200 papers

This study aims to enhance the quality of music generation using Transformers by incorporating meta-information. While Transformer-based approaches are effective at capturing long-term dependencies in musical compositions, the music they…

Sound · Computer Science 2026-05-21 Shinnosuke Taksuka , Hideo Mukai

Real-time music tracking systems follow a musical performance and at any time report the current position in a corresponding score. Most existing methods approach this problem exclusively in the audio domain, typically using online time…

Sound · Computer Science 2025-05-09 Silvan Peter , Patricia Hu , Gerhard Widmer

Beat tracking in musical performance MIDI is a challenging and important task for notation-level music transcription and rhythmical analysis, yet existing methods primarily focus on audio-based approaches. This paper proposes an end-to-end…

Sound · Computer Science 2025-07-02 Sebastian Murgul , Michael Heizmann

While end-to-end systems are becoming popular in auditory signal processing including automatic music tagging, models using raw audio as input needs a large amount of data and computational resources without domain knowledge. Inspired by…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-29 Yinghao Ma , Richard M. Stern

There has been fascinating work on creating artistic transformations of images by Gatys. This was revolutionary in how we can in some sense alter the 'style' of an image while generally preserving its 'content'. In our work, we present a…

Sound · Computer Science 2024-12-24 Prateek Verma , Julius O. Smith

Following the successful application of vision transformers in multiple computer vision tasks, these models have drawn the attention of the signal processing community. This is because signals are often represented as spectrograms (e.g.…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Nicolae-Catalin Ristea , Radu Tudor Ionescu , Fahad Shahbaz Khan

We study the merit of transfer learning for two sound recognition problems, i.e., audio tagging and sound event detection. Employing feature fusion, we adapt a baseline system utilizing only spectral acoustic inputs to also make use of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-27 Wim Boes , Hugo Van hamme

Time series analysis faces significant challenges in handling variable-length data and achieving robust generalization. While Transformer-based models have advanced time series tasks, they often struggle with feature redundancy and limited…

Machine Learning · Computer Science 2025-09-23 Kai Zhang , Siming Sun , Zhengyu Fan , Qinmin Yang , Xuejun Jiang

DJ mix transcription is a crucial step towards DJ mix reverse engineering, which estimates the set of parameters and audio effects applied to a set of existing tracks to produce a performative DJ mix. We introduce a new approach based on a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Étienne Paul André , Dominique Fourer , Diemo Schwarz

Guitar tablature is a form of music notation widely used among guitarists. It captures not only the musical content of a piece, but also its implementation and ornamentation on the instrument. Guitar Tablature Transcription (GTT) is an…

Sound · Computer Science 2025-05-27 Yongyi Zang , Yi Zhong , Frank Cwitkowitz , Zhiyao Duan

Little research focuses on cross-modal correlation learning where temporal structures of different data modalities such as audio and lyrics are taken into account. Stemming from the characteristic of temporal structures of music in nature,…

Information Retrieval · Computer Science 2017-11-30 Yi Yu , Suhua Tang , Francisco Raposo , Lei Chen

Tracking beats of singing voices without the presence of musical accompaniment can find many applications in music production, automatic song arrangement, and social media interaction. Its main challenge is the lack of strong rhythmic and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-01 Mojtaba Heydari , Zhiyao Duan

Multi-target tracking (MTT) is a classical signal processing task, where the goal is to estimate the states of an unknown number of moving targets from noisy sensor measurements. In this paper, we revisit MTT from a deep learning…

Signal Processing · Electrical Eng. & Systems 2024-05-15 Damian Owerko , Charilaos I. Kanatsoulis , Jennifer Bondarchuk , Donald J. Bucci , Alejandro Ribeiro

Transfer learning is critical for efficient information transfer across multiple related learning problems. A simple, yet effective transfer learning approach utilizes deep neural networks trained on a large-scale task for feature…

Sound · Computer Science 2021-06-23 Anurag Kumar , Yun Wang , Vamsi Krishna Ithapu , Christian Fuegen

The recent trend in multiple object tracking (MOT) is heading towards leveraging deep learning to boost the tracking performance. In this paper, we propose a novel solution named TransSTAM, which leverages Transformer to effectively model…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Peng Dai , Yiqiang Feng , Renliang Weng , Changshui Zhang

Rhythm patterns can be performed with a wide variation of tempi. This presents a challenge for many music information retrieval (MIR) systems; ideally, perceptually similar rhythms should be represented and processed similarly, regardless…

Sound · Computer Science 2018-05-01 Anders Elowsson

Capturing intricate and subtle variations in human expressiveness in music performance using computational approaches is challenging. In this paper, we propose a novel approach for reconstructing human expressiveness in piano performance…

Sound · Computer Science 2023-10-03 Jingjing Tang , Geraint Wiggins , Gyorgy Fazekas

Singing voice transcription converts recorded singing audio to musical notation. Sound contamination (such as accompaniment) and lack of annotated data make singing voice transcription an extremely difficult task. We take two approaches to…

Sound · Computer Science 2023-04-25 Xiangming Gu , Wei Zeng , Jianan Zhang , Longshen Ou , Ye Wang

Reconstructing the room transfer functions needed to calculate the complex sound field in a room has several important real-world applications. However, an unpractical number of microphones is often required. Recently, in addition to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Francesca Ronchini , Luca Comanducci , Mirco Pezzoli , Fabio Antonacci , Augusto Sarti

In line with the human capacity to perceive the world by simultaneously processing and integrating high-dimensional inputs from multiple modalities like vision and audio, we propose a novel model, MAiVAR-T (Multimodal Audio-Image to Video…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Muhammad Bilal Shaikh , Douglas Chai , Syed Mohammed Shamsul Islam , Naveed Akhtar