English
Related papers

Related papers: The match file format: Encoding Alignments between…

200 papers

We propose a content-based system for matching video and background music. The system aims to address the challenges in music recommendation for new users or new music give short-form videos. To this end, we propose a cross-modal framework…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Yi-Shan Lee , Wei-Cheng Tseng , Fu-En Wang , Min Sun

Automatic piano transcription models are typically evaluated using simple frame- or note-wise information retrieval (IR) metrics. Such benchmark metrics do not provide insights into the transcription quality of specific musical aspects such…

Sound · Computer Science 2024-10-10 Patricia Hu , Lukáš Samuel Marták , Carlos Cancino-Chacón , Gerhard Widmer

When songs are composed or performed, there is often an intent by the singer/songwriter of expressing feelings or emotions through it. For humans, matching the emotiveness in a musical composition or performance with the subjective…

Audio-to-score alignment aims at generating an accurate mapping between a performance audio and the score of a given piece. Standard alignment methods are based on Dynamic Time Warping (DTW) and employ handcrafted features. We explore the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-29 Ruchit Agrawal , Simon Dixon

Music is an expressive form of communication often used to convey emotion in scenarios where "words are not enough". Part of this information lies in the musical composition where well-defined language exists. However, a significant amount…

Sound · Computer Science 2017-08-14 Iman Malik , Carl Henrik Ek

Numerous studies in the field of music generation have demonstrated impressive performance, yet virtually no models are able to directly generate music to match accompanying videos. In this work, we develop a generative music AI framework,…

Sound · Computer Science 2024-06-03 Jaeyong Kang , Soujanya Poria , Dorien Herremans

Music exists in various modalities, such as score images, symbolic scores, MIDI, and audio. Translations between each modality are established as core tasks of music information retrieval, such as automatic music transcription…

Sound · Computer Science 2026-04-08 Jongmin Jung , Dongmin Kim , Sihun Lee , Seola Cho , Hyungjoon Soh , Irmak Bukey , Chris Donahue , Dasaem Jeong

Symbolic music datasets with matched scores and performances are essential for many music information retrieval (MIR) tasks. Yet, existing resources often cover a narrow range of composers, lack performance variety, omit note-level…

Sound · Computer Science 2026-05-08 Ilya Borovik

This paper describes an open-source Python framework for handling datasets for music processing tasks, built with the aim of improving the reproducibility of research projects in music computing and assessing the generalization abilities of…

Multimedia · Computer Science 2021-12-28 Federico Simonetta , Stavros Ntalampiras , Federico Avanzini

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronization. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment…

Computer Vision and Pattern Recognition · Computer Science 2020-11-05 Soo-Whan Chung , Joon Son Chung , Hong-Goo Kang

This paper proposes a novel Transformer-based model for music score infilling, to generate a music passage that fills in the gap between given past and future contexts. While existing infilling approaches can generate a passage that…

Sound · Computer Science 2022-10-07 Chih-Pin Tan , Alvin W. Y. Su , Yi-Hsuan Yang

Existing methods for expressive music performance rendering rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and…

Sound · Computer Science 2025-12-03 Hong-Jie You , Jie-Jing Shao , Xiao-Wen Yang , Lin-Han Jia , Lan-Zhe Guo , Yu-Feng Li

In this paper we present a new definition of the distortion matrix for a score following framework based on DTW. The proposal consists of arranging the score information in a sequence of note combinations and learning a spectral pattern for…

Sound · Computer Science 2019-05-30 A. J. Muñoz-Montoro , P. Vera-Candeas , D. Suarez-Dou , R. Cortina

Recent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow like motion feature representations, which exhibit limited…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Chuang Gan , Deng Huang , Hang Zhao , Joshua B. Tenenbaum , Antonio Torralba

Partitura is a lightweight Python package for handling symbolic musical information. It provides easy access to features commonly used in music information retrieval tasks, like note arrays (lists of timed pitched events) and 2D piano roll…

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source and target…

Multimedia · Computer Science 2023-11-27 Arjun Somayazulu , Changan Chen , Kristen Grauman

We study indeterminacies in realization of ornaments and how they can be incorporated in a stochastic performance model applicable for music information processing such as score-performance matching. We point out the importance of temporal…

Artificial Intelligence · Computer Science 2016-08-04 Eita Nakamura , Nobutaka Ono , Shigeki Sagayama , Kenji Watanabe

In this paper we present preliminary work examining the relationship between the formation of expectations and the realization of musical performances, paying particular attention to expressive tempo and dynamics. To compute features that…

Sound · Computer Science 2017-09-13 Carlos Cancino-Chacón , Maarten Grachten , David R. W. Sears , Gerhard Widmer

In competitive sports it is often very hard to quantify the performance. A player to score or overtake may depend on only millesimal of seconds or millimeters. In racquet sports like tennis, table tennis and squash many events will occur in…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-04 Katalin Hajdu-Szucs , Nora Fenyvesi , Jozsef Steger , Gabor Vattay

Mastering is an essential step in music production, but it is also a challenging task that has to go through the hands of experienced audio engineers, where they adjust tone, space, and volume of a song. Remastering follows the same…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-18 Junghyun Koo , Seungryeol Paik , Kyogu Lee