中文
相关论文

相关论文: Onsets and Frames: Dual-Objective Piano Transcript…

200 篇论文

Timbre and pitch are the two main perceptual properties of musical sounds. Depending on the target applications, we sometimes prefer to focus on one of them, while reducing the effect of the other. Researchers have managed to hand-craft…

声音 · 计算机科学 2018-11-09 Yun-Ning Hung , Yi-An Chen , Yi-Hsuan Yang

This paper addresses the matching of short music audio snippets to the corresponding pixel location in images of sheet music. A system is presented that simultaneously learns to read notes, listens to music and matches the currently played…

机器学习 · 计算机科学 2016-12-16 Matthias Dorfer , Andreas Arzt , Gerhard Widmer

In this work, we derive a generic overcomplete frame thresholding scheme based on risk minimization. Overcomplete frames being favored for analysis tasks such as classification, regression or anomaly detection, we provide a way to leverage…

音频与语音处理 · 电气工程与系统科学 2017-12-27 Romain Cosentino , Randall Balestriero , Richard Baraniuk , Ankit Patel

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Qiang Huang , Thomas Hain

Most of the current supervised automatic music transcription (AMT) models lack the ability to generalize. This means that they have trouble transcribing real-world music recordings from diverse musical genres that are not presented in the…

声音 · 计算机科学 2021-07-30 Kin Wai Cheuk , Dorien Herremans , Li Su

Multi-instrument Automatic Music Transcription (AMT), or the decoding of a musical recording into semantic musical content, is one of the holy grails of Music Information Retrieval. Current AMT approaches are restricted to piano and (some)…

声音 · 计算机科学 2022-04-29 Ben Maman , Amit H. Bermano

In this paper, we present a neural network approach for synchronizing audio recordings of human piano performances with their corresponding loosely aligned MIDI files. The task is addressed using a Convolutional Recurrent Neural Network…

声音 · 计算机科学 2025-06-30 Sebastian Murgul , Moritz Reiser , Michael Heizmann , Christoph Seibert

Automatic music transcription (AMT) has achieved remarkable progress for instruments such as the piano, largely due to the availability of large-scale, high-quality datasets. In contrast, violin AMT remains underexplored due to limited…

声音 · 计算机科学 2025-08-21 Yueh-Po Peng , Ting-Kang Wang , Li Su , Vincent K. M. Cheung

State-of-the-art end-to-end Optical Music Recognition (OMR) has, to date, primarily been carried out using monophonic transcription techniques to handle complex score layouts, such as polyphony, often by resorting to simplifications or…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Antonio Ríos-Vila , Jorge Calvo-Zaragoza , Thierry Paquet

Supervised and unsupervised homography estimation methods depend on image pairs tailored to specific modalities to achieve high accuracy. However, their performance deteriorates substantially when applied to unseen modalities. To address…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Jinkun You , Jiaxin Cheng , Jie Zhang , Yicong Zhou

Automatic Music Transcription (AMT) converts audio recordings into symbolic musical representations. Training deep neural networks (DNNs) for AMT typically requires strongly aligned training pairs with precise frame-level annotations. Since…

声音 · 计算机科学 2025-11-19 Jonathan Yaffe , Ben Maman , Meinard Müller , Amit H. Bermano

The automated creation of accurate musical notation from an expressive human performance is a fundamental task in computational musicology. To this end, we present an end-to-end deep learning approach that constructs detailed musical scores…

声音 · 计算机科学 2024-10-02 Tim Beyer , Angela Dai

Rethinking how to model polyphonic transcription formally, we frame it as a reinforcement learning task. Such a task formulation encompasses the notion of a musical agent and an environment containing an instrument as well as the sound…

声音 · 计算机科学 2018-05-30 Rainer Kelz , Gerhard Widmer

Quality of data plays an important role in most deep learning tasks. In the speech community, transcription of speech recording is indispensable. Since the transcription is usually generated artificially, automatically finding errors in…

计算与语言 · 计算机科学 2019-07-23 Xiaofei Wang , Jinyi Yang , Ruizhi Li , Samik Sadhu , Hynek Hermansky

Traditional methods to tackle many music information retrieval tasks typically follow a two-step architecture: feature engineering followed by a simple learning algorithm. In these "shallow" architectures, feature engineering and learning…

声音 · 计算机科学 2015-11-18 Peter Li , Jiyuan Qian , Tian Wang

Human visual attention is subjective and biased according to the personal preference of the viewer, however, current works of saliency detection are general and objective, without counting the factor of the observer. This will make the…

计算机视觉与模式识别 · 计算机科学 2018-02-23 Sikun Lin , Pan Hui

Music performance synthesis aims to synthesize a musical score into a natural performance. In this paper, we borrow recent advances in text-to-speech synthesis and present the Deep Performer -- a novel system for score-to-audio music…

声音 · 计算机科学 2022-02-22 Hao-Wen Dong , Cong Zhou , Taylor Berg-Kirkpatrick , Julian McAuley

Linking sheet music images to audio recordings remains a key problem for the development of efficient cross-modal music retrieval systems. One of the fundamental approaches toward this task is to learn a cross-modal embedding space via deep…

声音 · 计算机科学 2023-09-22 Luis Carvalho , Tobias Washüttl , Gerhard Widmer

Optical Music Recognition (OMR) has made significant progress since its inception, with various approaches now capable of accurately transcribing music scores into digital formats. Despite these advancements, most so-called end-to-end OMR…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Antonio Ríos-Vila , Jorge Calvo-Zaragoza , David Rizo , Thierry Paquet

Audio-to-score alignment aims at generating an accurate mapping between a performance audio and the score of a given piece. Standard alignment methods are based on Dynamic Time Warping (DTW) and employ handcrafted features, which cannot be…

声音 · 计算机科学 2020-11-17 Ruchit Agrawal , Simon Dixon