English
Related papers

Related papers: PBSCR: The Piano Bootleg Score Composer Recognitio…

200 papers

As one of the most fundamental techniques in multimodal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model…

Computer Vision and Pattern Recognition · Computer Science 2023-06-09 Shuo Yang , Zhaopan Xu , Kai Wang , Yang You , Hongxun Yao , Tongliang Liu , Min Xu

This paper addresses the problem of cross-modal musical piece identification and retrieval: finding the appropriate recording(s) from a database given a sheet music query, and vice versa, working directly with audio and scanned sheet music…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-27 Luis Carvalho , Gerhard Widmer

Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large language models have shown extraordinary potential in music,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Mingni Tang , Jiajia Li , Lu Yang , Zhiqiang Zhang , Jinghao Tian , Zuchao Li , Lefei Zhang , Ping Wang

Representation learning focused on disentangling the underlying factors of variation in given data has become an important area of research in machine learning. However, most of the studies in this area have relied on datasets from the…

Machine Learning · Computer Science 2020-07-31 Ashis Pati , Siddharth Gururani , Alexander Lerch

One of the key limitations of modern deep learning approaches lies in the amount of data required to train them. Humans, by contrast, can learn to recognize novel categories from just a few examples. Instrumental to this rapid learning…

Computer Vision and Pattern Recognition · Computer Science 2019-08-20 Pavel Tokmakov , Yu-Xiong Wang , Martial Hebert

We revisit the problems of pitch spelling and tonality guessing with a new algorithm for their joint estimation from a MIDI file including information about the measure boundaries. Our algorithm does not only identify a global key but also…

Sound · Computer Science 2024-02-19 Augustin Bouquillard , Florent Jacquemard

Paired image-text data with subtle variations in-between (e.g., people holding surfboards vs. people holding shovels) hold the promise of producing Vision-Language Models with proper compositional understanding. Synthesizing such training…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Haoxin Li , Boyang Li

Music arrangement generation is a subtask of automatic music generation, which involves reconstructing and re-conceptualizing a piece with new compositional techniques. Such a generation process inevitably requires reference from the…

Sound · Computer Science 2020-08-18 Ziyu Wang , Ke Chen , Junyan Jiang , Yiyi Zhang , Maoran Xu , Shuqi Dai , Xianbin Gu , Gus Xia

The identification of structural differences between a music performance and the score is a challenging yet integral step of audio-to-score alignment, an important subtask of music information retrieval. We present a novel method to detect…

Sound · Computer Science 2021-02-16 Ruchit Agrawal , Daniel Wolff , Simon Dixon

Portrait composition plays a central role in portrait aesthetics and visual communication, yet existing datasets and benchmarks mainly focus on coarse aesthetic scoring, generic image aesthetics, or unconstrained portrait generation. This…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Yuyang Sha , Zijie Lou , Youyun Tang , Xiaochao Qu , Zheng Qu , Ben Xia , Haoxiang Li , Ting Liu , Luoqi Liu

Cross-modal retrieval aims to align different modalities via semantic similarity. However, existing methods often assume that image-text pairs are perfectly aligned, overlooking Noisy Correspondences in real data. These misaligned pairs…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Zhuoyao Liu , Yang Liu , Wentao Feng , Shudong Huang

We introduce CLaMP: Contrastive Language-Music Pre-training, which learns cross-modal representations between natural language and symbolic music using a music encoder and a text encoder trained jointly with a contrastive loss. To pre-train…

Sound · Computer Science 2023-10-19 Shangda Wu , Dingyao Yu , Xu Tan , Maosong Sun

The goal of score following is to track a musical performance, usually in the form of audio, in a corresponding score representation. Established methods mainly rely on computer-readable scores in the form of MIDI or MusicXML and achieve…

Machine Learning · Computer Science 2019-10-17 Florian Henkel , Rainer Kelz , Gerhard Widmer

Artistic style has been studied for centuries, and recent advances in machine learning create new possibilities for understanding it computationally. However, ensuring that machine-learning models produce insights aligned with the interests…

Sound · Computer Science 2025-05-15 Huw Cheston , Reuben Bance , Peter M. C. Harrison

For many music analysis problems, we need to know the presence of instruments for each time frame in a multi-instrument musical piece. However, such a frame-level instrument recognition task remains difficult, mainly due to the lack of…

Sound · Computer Science 2019-02-19 Yun-Ning Hung , Yi-An Chen , Yi-Hsuan Yang

Current state-of-the-art deep learning systems for visual object recognition and detection use purely supervised training with regularization such as dropout to avoid overfitting. The performance depends critically on the amount of labeled…

Computer Vision and Pattern Recognition · Computer Science 2015-04-16 Scott Reed , Honglak Lee , Dragomir Anguelov , Christian Szegedy , Dumitru Erhan , Andrew Rabinovich

This paper proposes a deep convolutional neural network for performing note-level instrument assignment. Given a polyphonic multi-instrumental music signal along with its ground truth or predicted notes, the objective is to assign an…

Sound · Computer Science 2021-07-30 Carlos Lordelo , Emmanouil Benetos , Simon Dixon , Sven Ahlbäck

Most work on musical score models (a.k.a. musical language models) for music transcription has focused on describing the local sequential dependence of notes in musical scores and failed to capture their global repetitive structure, which…

Sound · Computer Science 2021-02-17 Eita Nakamura , Kazuyoshi Yoshii

Pretraining robust vision or multimodal foundation models (e.g., CLIP) relies on large-scale datasets that may be noisy, potentially misaligned, and have long-tail distributions. Previous works have shown promising results in augmenting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Qingqing Cao , Mahyar Najibi , Sachin Mehta

In this paper, we explore the tokenized representation of musical scores using the Transformer model to automatically generate musical scores. Thus far, sequence models have yielded fruitful results with note-level (MIDI-equivalent)…

Sound · Computer Science 2021-12-02 Masahiro Suzuki