English
Related papers

Related papers: The GigaMIDI Dataset with Features for Expressive …

200 papers

Quantification of stylistic differences between musical artists is of academic interest to the music community, and is also useful for other applications such as music information retrieval and recommendation systems. Information about…

Applications · Statistics 2020-12-23 Anna K. Yanchenko , Peter D. Hoff

Recent MIDI-to-audio synthesis methods using deep neural networks have successfully generated high-quality, expressive instrumental tracks. However, these methods require MIDI annotations for supervised training, limiting the diversity of…

Sound · Computer Science 2025-06-12 Osamu Take , Taketo Akama

We present the MIDInfinite, a web application capable of generating symbolic music using a large-scale generative AI model locally on commodity hardware. Creating this demo involved porting the Anticipatory Music Transformer, a large…

Sound · Computer Science 2024-11-15 Xun Zhou , Charlie Ruan , Zihe Zhao , Tianqi Chen , Chris Donahue

Music scores are written representations of music and contain rich information about musical components. The visual information on music scores includes notes, rests, staff lines, clefs, dynamics, and articulations. This visual information…

Multimedia · Computer Science 2024-06-18 Yuheng Lin , Zheqi Dai , Qiuqiang Kong

In this paper, we introduce the Extreme Metal Vocals Dataset, which comprises a collection of recordings of extreme vocal techniques performed within the realm of heavy metal music. The dataset consists of 760 audio excerpts of 1 second to…

Sound · Computer Science 2024-06-26 Modan Tailleur , Julien Pinquier , Laurent Millot , Corsin Vogel , Mathieu Lagrange

Automatic Chord Estimation (ACE) is a fundamental task in Music Information Retrieval (MIR) and has applications in both music performance and MIR research. The task consists of segmenting a music recording or score and assigning a chord…

Sound · Computer Science 2020-02-25 Daphne Odekerken , Hendrik Vincent Koops , Anja Volk

This study focuses on the perception of music performances when contextual factors, such as room acoustics and instrument, change. We propose to distinguish the concept of "performance" from the one of "interpretation", which expresses the…

Sound · Computer Science 2022-03-08 Federico Simonetta , Federico Avanzini , Stavros Ntalampiras

Modern music retrieval systems often rely on fixed representations of user preferences, limiting their ability to capture users' diverse and uncertain retrieval needs. To address this limitation, we introduce Diff4Steer, a novel generative…

Sound · Computer Science 2025-04-25 Xuchan Bao , Judith Yue Li , Zhong Yi Wan , Kun Su , Timo Denk , Joonseok Lee , Dima Kuzmin , Fei Sha

Manual melody detection is a tedious task requiring high expertise level, while automatic detection is often not expressive or powerful enough. Thus, we present MelodyVis, a visual application designed in collaboration with musicology…

Human-Computer Interaction · Computer Science 2024-07-09 Matthias Miller , Daniel Fürst , Maximilian T. Fischer , Hanna Hauptmann , Daniel Keim , Mennatallah El-Assady

This work present a music dataset named MusicTM-Dataset, which is utilized in improving the representation learning ability of different types of cross-modal retrieval (CMR). Little large music dataset including three modalities is…

Sound · Computer Science 2021-05-10 Donghuo Zeng , Yi Yu , Keizo Oyama

This paper introduces the jazznet Dataset, a dataset of fundamental jazz piano music patterns for developing machine learning (ML) algorithms in music information retrieval (MIR). The dataset contains 162520 labeled piano patterns,…

Sound · Computer Science 2023-02-20 Tosiron Adegbija

Music Emotion Recogniser (MER) research faces challenges due to limited high-quality annotated datasets and difficulties in addressing cross-track feature drift. This work presents two primary contributions to address these issues.…

Sound · Computer Science 2025-12-18 Qilin Li , C. L. Philip Chen , Tong Zhang

This paper investigates a cross-modal retrieval problem in which a user would like to retrieve a passage of music from a MIDI file by taking a cell phone picture of a physical page of sheet music. While audio-sheet music retrieval has been…

Multimedia · Computer Science 2020-04-23 Daniel Yang , Thitaree Tanprasert , Teerapat Jenrungrot , Mengyi Shan , TJ Tsai

In the past, the field of drum source separation faced significant challenges due to limited data availability, hindering the adoption of cutting-edge deep learning methods that have found success in other related audio applications. In…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-21 Alessandro Ilic Mezza , Riccardo Giampiccolo , Alberto Bernardini , Augusto Sarti

Moonbeam is a transformer-based foundation model for symbolic music, pretrained on a large and diverse collection of MIDI data totaling 81.6K hours of music and 18 billion tokens. Moonbeam incorporates music-domain inductive biases by…

Sound · Computer Science 2025-05-22 Zixun Guo , Simon Dixon

This thesis combines audio-analysis with computer vision to approach Music Information Retrieval (MIR) tasks from a multi-modal perspective. This thesis focuses on the information provided by the visual layer of music videos and how it can…

Multimedia · Computer Science 2020-02-04 Alexander Schindler

In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments. We first identify two key intermediate representations for a successful video to music…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Chuang Gan , Deng Huang , Peihao Chen , Joshua B. Tenenbaum , Antonio Torralba

We consider the task of multimodal music mood prediction based on the audio signal and the lyrics of a track. We reproduce the implementation of traditional feature engineering based approaches and propose a new model based on deep…

Information Retrieval · Computer Science 2018-09-21 Rémi Delbouys , Romain Hennequin , Francesco Piccoli , Jimena Royo-Letelier , Manuel Moussallam

In this work, we introduce the Sheet Music Benchmark (SMB), a dataset of six hundred and eighty-five pages specifically designed to benchmark Optical Music Recognition (OMR) research. SMB encompasses a diverse array of musical textures,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Juan C. Martinez-Sevilla , Joan Cerveto-Serrano , Noelia Luna , Greg Chapman , Craig Sapp , David Rizo , Jorge Calvo-Zaragoza

Music representation learning is central to music information retrieval and generation. While recent advances in multimodal learning have improved alignment between text and audio for tasks such as cross-modal music retrieval, text-to-music…