English
Related papers

Related papers: Sheet Music Benchmark: Standardized Optical Music …

200 papers

Deep learning has recently been applied to optical music recognition (OMR). However, currently OMR processing from various sheet music images still lacks precision to be widely applicable. Here, we present an MMdA (Measure-based Multimodal…

Computer Vision and Pattern Recognition · Computer Science 2021-06-24 Tomoyuki Shishido , Fehmiju Fati , Daisuke Tokushige , Yasuhiro Ono

Enhancing the ability of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) to interpret sheet music is a crucial step toward building AI musicians. However, current research lacks both evaluation benchmarks and…

Computation and Language · Computer Science 2025-09-29 Zhilin Wang , Zhe Yang , Yun Luo , Yafu Li , Xiaoye Qu , Ziqian Qiao , Haoran Zhang , Runzhe Zhan , Derek F. Wong , Jizhe Zhou , Yu Cheng

Music Emotion Recognition (MER) is a task deeply connected to human perception, relying heavily on subjective annotations collected from contributors. Prior studies tend to focus on specific musical styles rather than incorporating a…

Sound · Computer Science 2025-11-14 Joann Ching , Gerhard Widmer

Multimodal music emotion recognition (MMER) is an emerging discipline in music information retrieval that has experienced a surge in interest in recent years. This survey provides a comprehensive overview of the current state-of-the-art in…

Multimedia · Computer Science 2025-04-29 Rashini Liyanarachchi , Aditya Joshi , Erik Meijering

Music Source Restoration (MSR) extends source separation to realistic settings where signals undergo production effects (equalization, compression, reverb) and real-world degradations, with the goal of recovering the original unprocessed…

This work addresses the problem of matching short excerpts of audio with their respective counterparts in sheet music images. We show how to employ neural network-based cross-modality embedding spaces for solving the following two sheet…

Information Retrieval · Computer Science 2017-08-01 Matthias Dorfer , Andreas Arzt , Gerhard Widmer

Symbolic Music Emotion Recognition(SMER) is to predict music emotion from symbolic data, such as MIDI and MusicXML. Previous work mainly focused on learning better representation via (mask) language model pre-training but ignored the…

Sound · Computer Science 2022-01-19 Jibao Qiu , C. L. Philip Chen , Tong Zhang

Optical music recognition (OMR) aims to convert music notation into digital formats. One approach to tackle OMR is through a multi-stage pipeline, where the system first detects visual music notation elements in the image (object detection)…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Guang Yang , Muru Zhang , Lin Qiu , Yanming Wan , Noah A. Smith

Music structure analysis (MSA) methods traditionally search for musically meaningful patterns in audio: homogeneity, repetition, novelty, and segment-length regularity. Hand-crafted audio features such as MFCCs or chromagrams are often used…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-03 Ju-Chiang Wang , Jordan B. L. Smith , Wei-Tsung Lu , Xuchen Song

We propose Legato, a new end-to-end model for optical music recognition (OMR), a task of converting music score images to machine-readable documents. Legato is the first large-scale pretrained OMR model capable of recognizing full-page or…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Guang Yang , Victoria Ebert , Nazif Tamer , Brian Siyuan Zheng , Luiza Pozzobon , Noah A. Smith

This work present a music dataset named MusicTM-Dataset, which is utilized in improving the representation learning ability of different types of cross-modal retrieval (CMR). Little large music dataset including three modalities is…

Sound · Computer Science 2021-05-10 Donghuo Zeng , Yi Yu , Keizo Oyama

Optical Music Recognition (OMR) is an important and challenging area within music information retrieval, the accurate detection of music symbols in digital images is a core functionality of any OMR pipeline. In this paper, we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2018-05-29 Lukas Tuggener , Ismail Elezi , Jurgen Schmidhuber , Thilo Stadelmann

This paper addresses the task of score following in sheet music given as unprocessed images. While existing work either relies on OMR software to obtain a computer-readable score representation, or crucially relies on prepared sheet image…

Machine Learning · Computer Science 2020-07-22 Florian Henkel , Rainer Kelz , Gerhard Widmer

In the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the…

Learning from noisy labels remains a major challenge in medical image analysis, where annotation demands expert knowledge and substantial inter-observer variability often leads to inconsistent or erroneous labels. Despite extensive research…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yuan Ma , Junlin Hou , Chao Zhang , Yukun Zhou , Zongyuan Ge , Haoran Xie , Lie Ju

We address the problem of detecting the number of complex exponentials and estimating their parameters from a noisy signal using the Matrix Pencil (MP) method. We introduce the MP modes and present their informative spectral structure. We…

Signal Processing · Electrical Eng. & Systems 2025-09-30 Yehonatan-Itay Segman , Alon Amar , Ronen Talmon

In this paper, we bridge the gap between visualization and musicology by focusing on rhythm analysis tasks, which are tedious due to the complex visual encoding of the well-established Common Music Notation (CMN). Instead of replacing the…

Human-Computer Interaction · Computer Science 2020-09-07 Daniel Fürst , Matthias Miller , Daniel Keim , Alexandra Bonnici , Hanna Schäfer , Mennatallah El-Assady

Music source separation (MSS) is a task that involves isolating individual sound sources, or stems, from mixed audio signals. This paper presents an ensemble approach to MSS, combining several state-of-the-art architectures to achieve…

Sound · Computer Science 2024-10-29 Saarth Vardhan , Pavani R Acharya , Samarth S Rao , Oorjitha Ratna Jasthi , S Natarajan

We present a robust refinement method for estimating oriented normals from unstructured point clouds. In contrast to previous approaches that either suffer from high computational complexity or fail to achieve desirable accuracy, our novel…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Yingrui Wu , Mingyang Zhao , Weize Quan , Jian Shi , Xiaohong Jia , Dong-Ming Yan

We propose a framework for audio-to-score alignment on piano performance that employs automatic music transcription (AMT) using neural networks. Even though the AMT result may contain some errors, the note prediction output can be regarded…

Sound · Computer Science 2017-11-15 Taegyun Kwon , Dasaem Jeong , Juhan Nam