中文
相关论文

相关论文: Pay Attention to the Keys: Visual Piano Transcript…

200 篇论文

Most recent research about automatic music transcription (AMT) uses convolutional neural networks and recurrent neural networks to model the mapping from music signals to symbolic notation. Based on a high-resolution piano transcription…

音频与语音处理 · 电气工程与系统科学 2022-04-11 Longshen Ou , Ziyi Guo , Emmanouil Benetos , Jiqing Han , Ye Wang

Automatic music transcription (AMT) is the task of transcribing audio recordings into symbolic representations. Recently, neural network-based methods have been applied to AMT, and have achieved state-of-the-art results. However, many…

声音 · 计算机科学 2021-08-03 Qiuqiang Kong , Bochen Li , Xuchen Song , Yuan Wan , Yuxuan Wang

While automatic music transcription is well-established in music information retrieval, most models are limited to transcribing pitch and timing information from audio, and thus omit crucial expressive and instrument-specific nuances. One…

声音 · 计算机科学 2026-02-04 Ting-Kang Wang , Yueh-Po Peng , Li Su , Vincent K. M. Cheung

We advance the state of the art in polyphonic piano music transcription by using a deep convolutional and recurrent neural network which is trained to jointly predict onsets and frames. Our model predicts pitch onset events and then uses…

Many social media users prefer consuming content in the form of videos rather than text. However, in order for content creators to produce videos with a high click-through rate, much editing is needed to match the footage to the music. This…

机器学习 · 计算机科学 2022-01-03 Chin-Tung Lin , Mu Yang

This paper describes a streaming audio-to-MIDI piano transcription approach that aims to sequentially translate a music signal into a sequence of note onset and offset events. The sequence-to-sequence nature of this task may call for the…

声音 · 计算机科学 2025-03-04 Weixing Wei , Jiahao Zhao , Yulun Wu , Kazuyoshi Yoshii

Polyphonic Piano Transcription has recently experienced substantial progress, driven by the use of sophisticated Deep Learning approaches and the introduction of new subtasks such as note onset, offset, velocity and pedal detection. This…

声音 · 计算机科学 2023-06-02 Andres Fernandez

Many of the recent approaches to polyphonic piano note onset transcription require training a machine learning model on a large piano database. However, such approaches are limited by dataset availability; additional training data is…

机器学习 · 统计学 2017-07-27 Samuel Li

Automatic music transcription (AMT) has achieved remarkable progress for instruments such as the piano, largely due to the availability of large-scale, high-quality datasets. In contrast, violin AMT remains underexplored due to limited…

声音 · 计算机科学 2025-08-21 Yueh-Po Peng , Ting-Kang Wang , Li Su , Vincent K. M. Cheung

This paper describes a novel paradigm that formalizes automatic piano transcription (APT) as an optimal transport (OT) problem, not as a frame-level multi-label binary classification problem. Our method learns to minimize the cost of…

声音 · 计算机科学 2026-05-19 Weixing Wei , Raynaldi Lalang , Dichucheng Li , Kazuyoshi Yoshii

Taking long-term spectral and temporal dependencies into account is essential for automatic piano transcription. This is especially helpful when determining the precise onset and offset for each note in the polyphonic piano content. In this…

声音 · 计算机科学 2023-07-11 Keisuke Toyama , Taketo Akama , Yukara Ikemiya , Yuhta Takida , Wei-Hsiang Liao , Yuki Mitsufuji

This paper presents a statistical method for use in music transcription that can estimate score times of note onsets and offsets from polyphonic MIDI performance signals. Because performed note durations can deviate largely from…

人工智能 · 计算机科学 2017-07-10 Eita Nakamura , Kazuyoshi Yoshii , Simon Dixon

While piano music transcription models have shown high performance for solo piano recordings, their performance degrades when applied to ensemble recordings. This study aims to analyze the impact of different data augmentation methods on…

声音 · 计算机科学 2023-05-24 Hyemi Kim , Jiyun Park , Taegyun Kwon , Dasaem Jeong , Juhan Nam

We propose a framework for audio-to-score alignment on piano performance that employs automatic music transcription (AMT) using neural networks. Even though the AMT result may contain some errors, the note prediction output can be regarded…

声音 · 计算机科学 2017-11-15 Taegyun Kwon , Dasaem Jeong , Juhan Nam

Algorithms for automatic piano transcription have improved dramatically in recent years due to new datasets and modeling techniques. Recent developments have focused primarily on adapting new neural network architectures, such as the…

声音 · 计算机科学 2024-02-05 Drew Edwards , Simon Dixon , Emmanouil Benetos , Akira Maezawa , Yuta Kusaka

While neural network models are making significant progress in piano transcription, they are becoming more resource-consuming due to requiring larger model size and more computing power. In this paper, we attempt to apply more prior about…

声音 · 计算机科学 2022-09-01 Weixing Wei , Peilin Li , Yi Yu , Wei Li

Motivated by the state-of-art psychological research, we note that a piano performance transcribed with existing Automatic Music Transcription (AMT) methods cannot be successfully resynthesized without affecting the artistic content of the…

声音 · 计算机科学 2026-01-21 Federico Simonetta , Stavros Ntalampiras , Federico Avanzini

Viewing polyphonic piano transcription as a multitask learning problem, where we need to simultaneously predict onsets, intermediate frames and offsets of notes, we investigate the performance impact of additional prediction targets, using…

声音 · 计算机科学 2019-02-13 Rainer Kelz , Sebastian Böck , Gerhard Widmer

Automatic Music Transcription (AMT), aiming to get musical notes from raw audio, typically uses frame-level systems with piano-roll outputs or language model (LM)-based systems with note-level predictions. However, frame-level systems…

声音 · 计算机科学 2025-01-08 Dichucheng Li , Yongyi Zang , Qiuqiang Kong

Acoustic events often have a visual counterpart. Knowledge of visual information can aid the understanding of complex auditory scenes, even when only a stereo mixdown is available in the audio domain, \eg identifying which musicians are…

神经与进化计算 · 计算机科学 2017-06-30 A. Bazzica , J. C. van Gemert , C. C. S. Liem , A. Hanjalic
‹ 上一页 1 2 3 10 下一页 ›