English
Related papers

Related papers: Discogs-VI: A Musical Version Identification Datas…

200 papers

Version identification (VI) systems now offer accurate and scalable solutions for detecting different renditions of a musical composition, allowing the use of these systems in industrial applications and throughout the wider music…

Sound · Computer Science 2021-10-01 Furkan Yesiler , Marius Miron , Joan Serrà , Emilia Gómez

In this article, we aim to provide a review of the key ideas and approaches proposed in 20 years of scientific literature around musical version identification (VI) research and connect them to current practice. For more than a decade, VI…

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition models. We make three…

Computer Vision and Pattern Recognition · Computer Science 2020-09-28 Honglie Chen , Weidi Xie , Andrea Vedaldi , Andrew Zisserman

Version identification (VI) has seen substantial progress over the past few years. On the one hand, the introduction of the metric learning paradigm has favored the emergence of scalable yet accurate VI systems. On the other hand, using…

Sound · Computer Science 2022-10-05 Mathilde Abrassart , Guillaume Doras

The setlist identification (SLI) task addresses a music recognition use case where the goal is to retrieve the metadata and timestamps for all the tracks played in live music events. Due to various musical and non-musical changes in live…

Sound · Computer Science 2021-01-07 Furkan Yesiler , Emilio Molina , Joan Serrà , Emilia Gómez

We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural language descriptions…

The goal of this paper is twofold. First, we introduce DALI, a large and rich multimodal dataset containing 5358 audio tracks with their time-aligned vocal melody notes and lyrics at four levels of granularity. The second goal is to explain…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-26 Gabriel Meseguer-Brocal , Alice Cohen-Hadria , Geoffroy Peeters

A large-scale dataset is essential for training a well-generalized deep-learning model. Most such datasets are collected via scraping from various internet sources, inevitably introducing duplicated data. In the symbolic music domain, these…

Sound · Computer Science 2025-09-23 Eunjin Choi , Hyerin Kim , Jiwoo Ryu , Juhan Nam , Dasaem Jeong

Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist due to the lack of a captioned MIDI dataset. This work aims to enable research that combines LLMs with symbolic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-08 Jan Melechovsky , Abhinaba Roy , Dorien Herremans

Recent advances in AI-based music generation have focused heavily on text-conditioned models, with less attention given to reference-based generation such as song adaptation. To support this line of research, we introduce LargeSHS, a…

Sound · Computer Science 2025-11-25 Chih-Pin Tan , Hsuan-Kai Kao , Li Su , Yi-Hsuan Yang

In this document, we introduce a new dataset designed for training machine learning models of symbolic music data. Five datasets are provided, one of which is from a newly collected corpus of 20K midi files. We describe our preprocessing…

Sound · Computer Science 2016-06-09 Christian Walder

The version identification (VI) task deals with the automatic detection of recordings that correspond to the same underlying musical piece. Despite many efforts, VI is still an open problem, with much room for improvement, specially with…

Sound · Computer Science 2020-04-14 Furkan Yesiler , Joan Serrà , Emilia Gómez

Neural models are one of the most popular approaches for music generation, yet there aren't standard large datasets tailored for learning music directly from game data. To address this research gap, we introduce a novel dataset named…

Sound · Computer Science 2024-04-09 Igor Cardoso , Rubens O. Moraes , Lucas N. Ferreira

Musical dynamics form a core part of expressive singing voice performances. However, automatic analysis of musical dynamics for singing voice has received limited attention partly due to the scarcity of suitable datasets and a lack of clear…

Sound · Computer Science 2024-10-29 Jyoti Narang , Nazif Can Tamer , Viviana De La Vega , Xavier Serra

Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large language models have shown extraordinary potential in music,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Mingni Tang , Jiajia Li , Lu Yang , Zhiqiang Zhang , Jinghao Tian , Zuchao Li , Lefei Zhang , Ping Wang

Today, comprehensive evaluation of large-scale machine learning models is possible thanks to the open datasets produced using crowdsourcing, such as SQuAD, MS COCO, ImageNet, SuperGLUE, etc. These datasets capture objective responses,…

Human-Computer Interaction · Computer Science 2021-11-29 Nikita Pavlichenko , Dmitry Ustalov

We present a new large-scale emotion-labeled symbolic music dataset consisting of 12k MIDI songs. To create this dataset, we first trained emotion classification models on the GoEmotions dataset, achieving state-of-the-art results with a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-28 Serkan Sulun , Pedro Oliveira , Paula Viana

Audio classification is the task of identifying the sound categories that are associated with a given audio signal. This paper presents an investigation on large-scale audio classification based on the recently released AudioSet database.…

Sound · Computer Science 2018-10-31 Yuzhong Wu , Tan Lee

At the core of many important machine learning problems faced by online streaming services is a need to model how users interact with the content they are served. Unfortunately, there are no public datasets currently available that enable…

Information Retrieval · Computer Science 2020-10-16 Brian Brost , Rishabh Mehrotra , Tristan Jehan

One of the main limitations in the field of audio signal processing is the lack of large public datasets with audio representations and high-quality annotations due to restrictions of copyrighted commercial music. We present Melon Playlist…

‹ Prev 1 2 3 10 Next ›