中文
相关论文

相关论文: Evaluating Non-aligned Musical Score Transcription…

200 篇论文

Understanding complete musical scores entails integrated reasoning over pitch, rhythm, harmony, and large-scale structure, yet the ability of Large Language Models and Vision--Language Models to interpret full musical notation remains…

Audio alignment is a fundamental preprocessing step in many MIR pipelines. For two audio clips with M and N frames, respectively, the most popular approach, dynamic time warping (DTW), has O(MN) requirements in both memory and computation,…

声音 · 计算机科学 2020-08-07 Christopher Tralie , Elizabeth Dempsey

Forced alignment refers to a technology that time-aligns a given transcription with a corresponding speech. However, as the forced alignment technologies have developed using speech audio, they might fail in alignment when the input speech…

计算机视觉与模式识别 · 计算机科学 2023-03-16 Minsu Kim , Chae Won Kim , Yong Man Ro

Automatic Music Transcription (AMT), inferring musical notes from raw audio, is a challenging task at the core of music understanding. Unlike Automatic Speech Recognition (ASR), which typically focuses on the words of a single speaker, AMT…

声音 · 计算机科学 2022-03-16 Josh Gardner , Ian Simon , Ethan Manilow , Curtis Hawthorne , Jesse Engel

Automatic Music Transcription (AMT) converts audio recordings into symbolic musical representations. Training deep neural networks (DNNs) for AMT typically requires strongly aligned training pairs with precise frame-level annotations. Since…

声音 · 计算机科学 2025-11-19 Jonathan Yaffe , Ben Maman , Meinard Müller , Amit H. Bermano

Multi-instrument music transcription aims to convert polyphonic music recordings into musical scores assigned to each instrument. This task is challenging for modeling as it requires simultaneously identifying multiple instruments and…

音频与语音处理 · 电气工程与系统科学 2024-08-02 Sungkyun Chang , Emmanouil Benetos , Holger Kirchhoff , Simon Dixon

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source and target…

多媒体 · 计算机科学 2023-11-27 Arjun Somayazulu , Changan Chen , Kristen Grauman

Real-time computer-based accompaniment for human musical performances entails three critical tasks: identifying what the performer is playing, locating their position within the score, and synchronously playing the accompanying parts. Among…

声音 · 计算机科学 2025-03-11 Ashwin Pillay

Lyrics transcription of polyphonic music is challenging because singing vocals are corrupted by the background music. To improve the robustness of lyrics transcription to the background music, we propose a strategy of combining the features…

音频与语音处理 · 电气工程与系统科学 2022-04-25 Xiaoxue Gao , Chitralekha Gupta , Haizhou Li

We present a probabilistic generative model for timing deviations in expressive music performance. The structure of the proposed model is equivalent to a switching state space model. The switch variables correspond to discrete note…

人工智能 · 计算机科学 2011-06-27 A. T. Cemgil , B. Kappen

Music scores are used to precisely store music pieces for transmission and preservation. To represent and manipulate these complex objects, various formats have been tailored for different use cases. While music notation follows specific…

多媒体 · 计算机科学 2025-10-06 Géré Léo , Nicolas Audebert , Florent Jacquemard

Generating full-length, high-quality songs is challenging, as it requires maintaining long-term coherence both across text and music modalities and within the music modality itself. Existing non-autoregressive (NAR) frameworks, while…

音频与语音处理 · 电气工程与系统科学 2026-02-04 Yuepeng Jiang , Huakang Chen , Ziqian Ning , Jixun Yao , Zerui Han , Di Wu , Meng Meng , Jian Luan , Zhonghua Fu , Lei Xie

In this paper, we explore the tokenized representation of musical scores using the Transformer model to automatically generate musical scores. Thus far, sequence models have yielded fruitful results with note-level (MIDI-equivalent)…

声音 · 计算机科学 2021-12-02 Masahiro Suzuki

We propose a novel method to model hierarchical metrical structures for both symbolic music and audio signals in a self-supervised manner with minimal domain knowledge. The model trains and inferences on beat-aligned music signals and…

声音 · 计算机科学 2023-01-26 Junyan Jiang , Gus Xia

Score following is the process of tracking a musical performance (audio) with respect to a known symbolic representation (a score). We start this paper by formulating score following as a multimodal Markov Decision Process, the mathematical…

人工智能 · 计算机科学 2018-07-18 Matthias Dorfer , Florian Henkel , Gerhard Widmer

As a result of continuous advances in Music Information Retrieval (MIR) technology, generating and distributing music has become more diverse and accessible. In this context, interest in music intellectual property protection is increasing…

人工智能 · 计算机科学 2026-02-03 Seonghyeon Go

A fitting soundtrack can help a video better convey its content and provide a better immersive experience. This paper introduces a novel approach utilizing self-supervised learning and contrastive learning to automatically recommend audio…

多媒体 · 计算机科学 2025-03-10 Shimiao Liu , Alexander Lerch

Modelling human perception of musical similarity is critical for the evaluation of generative music systems, musicological research, and many Music Information Retrieval tasks. Although human similarity judgments are the gold standard,…

音频与语音处理 · 电气工程与系统科学 2020-06-29 Jeff Ens , Philippe Pasquier

Automatic lyrics to polyphonic audio alignment is a challenging task not only because the vocals are corrupted by background music, but also there is a lack of annotated polyphonic corpus for effective acoustic modeling. In this work, we…

音频与语音处理 · 电气工程与系统科学 2019-06-26 Chitralekha Gupta , Emre Yılmaz , Haizhou Li

The paper presents methods for evaluating the accuracy of alignments between transcriptions and audio recordings. The methods have been applied to the Spoken British National Corpus, which is an extensive and varied corpus of natural…

声音 · 计算机科学 2011-01-11 Ladan Baghai-Ravary , Sergio Grau , Greg Kochanski