English
Related papers

Related papers: Playing Technique Detection by Fusing Note Onset I…

200 papers

Automatic music transcription (AMT) aims to infer a latent symbolic representation of a piece of music (piano-roll), given a corresponding observed audio recording. Transcribing polyphonic music (when multiple notes are played…

Machine Learning · Statistics 2018-11-19 Pablo A. Alvarado , Dan Stowell

Visual piano transcription (VPT) is the task of obtaining a symbolic representation of a piano performance from visual information alone (e.g., from a top-down video of the piano keyboard). In this work we propose a VPT system based on the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Uros Zivanovic , Ivan Pilkov , Carlos Eduardo Cancino-Chacón

Estimating the performance difficulty of a musical score is crucial in music education for adequately designing the learning curriculum of the students. Although the Music Information Retrieval community has recently shown interest in this…

Sound · Computer Science 2023-09-29 Pedro Ramoneda , Jose J. Valero-Mas , Dasaem Jeong , Xavier Serra

MIDI-sheet music alignment is the task of finding correspondences between a MIDI representation of a piece and its corresponding sheet music images. Rather than using optical music recognition to bridge the gap between sheet music and MIDI,…

Multimedia · Computer Science 2020-04-23 Thitaree Tanprasert , Teerapat Jenrungrot , Meinard Mueller , T. J. Tsai

In recent years, the task of Automatic Music Transcription (AMT), whereby various attributes of music notes are estimated from audio, has received increasing attention. At the same time, the related task of Multi-Pitch Estimation (MPE)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-18 Frank Cwitkowitz , Toni Hirvonen , Anssi Klapuri

This paper proposes a robust ear identification system which is developed by fusing SIFT features of color segmented slice regions of an ear. The proposed ear identification method makes use of Gaussian mixture model (GMM) to build ear…

Computer Vision and Pattern Recognition · Computer Science 2010-07-23 Dakshina Ranjan Kisku , Phalguni Gupta , Jamuna Kanta Sing

Prompt-based methods have achieved promising results in most few-shot text classification tasks. However, for readability assessment tasks, traditional prompt methods lackcrucial linguistic knowledge, which has already been proven to be…

Computation and Language · Computer Science 2024-04-11 Ziyang Wang , Sanwoo Lee , Hsiu-Yuan Huang , Yunfang Wu

The clustering performance of Fuzzy Adaptive Resonance Theory (Fuzzy ART) is highly dependent on the preset vigilance parameter, where deviations in its value can lead to significant fluctuations in clustering results, severely limiting its…

Machine Learning · Computer Science 2025-05-09 Xiaozheng Qu , Zhaochuan Li , Zhuang Qi , Xiang Li , Haibei Huang , Lei Meng , Xiangxu Meng

Generative Pre-trained Transformer models, known as GPT or OPT, set themselves apart through breakthrough performance across complex language modelling tasks, but also by their extremely high computational and storage costs. Specifically,…

Machine Learning · Computer Science 2023-03-23 Elias Frantar , Saleh Ashkboos , Torsten Hoefler , Dan Alistarh

Intensity estimation for Poisson processes is a classical problem and has been extensively studied over the past few decades. Practical observations, however, often contain compositional noise, i.e. a nonlinear shift along the time axis,…

Methodology · Statistics 2019-09-25 Glenna Schluck , Wei Wu , Anuj Srivastava

Disfluencies in spontaneous speech are known to be associated with prosodic disruptions. However, most algorithms for disfluency detection use only word transcripts. Integrating prosodic cues has proved difficult because of the many sources…

Computation and Language · Computer Science 2019-04-10 Vicky Zayats , Mari Ostendorf

Onsets are a key factor to split audio into several notes. In this paper, we ensemble multiple temporal convolution network (TCN) based model and utilize a restricted frequency range spectrogram to achieve more robust onset detection.…

Sound · Computer Science 2023-06-09 Yu Cheng Hung , Jian-Jiun Ding

This article presents an Integral Gaussian Process (IntegralGP) framework for volumetric estimation of subterranean properties in mineral deposits. It provides a unified representation for data with different spatial supports, which enables…

Methodology · Statistics 2025-12-11 Anna Chlingaryan , Arman Melkumyan , Raymond Leung

Musical instruments recognition is a widely used application for music information retrieval. As most of previous musical instruments recognition dataset focus on western musical instruments, it is difficult for researcher to study and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-14 Xia Gong , Yuxiang Zhu , Haidi Zhu , Haoran Wei

Intent classification is a task in spoken language understanding. An intent classification system is usually implemented as a pipeline process, with a speech recognition module followed by text processing that classifies the intents. There…

Computation and Language · Computer Science 2021-02-16 Bidisha Sharma , Maulik Madhavi , Haizhou Li

Ensemble learning for anomaly detection of data structured into complex network has been barely studied due to the inconsistent performance of complex network characteristics and lack of inherent objective function. In this paper, we…

Social and Information Networks · Computer Science 2018-07-25 Jinfa Wang , Xiao Liu , Hai Zhao , Xingchi Chen

Vocal Percussion Transcription (VPT) is concerned with the automatic detection and classification of vocal percussion sound events, allowing music creators and producers to sketch drum lines on the fly. Classifier algorithms in VPT systems…

Sound · Computer Science 2022-04-12 Alejandro Delgado , Emir Demirel , Vinod Subramanian , Charalampos Saitis , Mark Sandler

Benefiting from prompt tuning, recent years have witnessed the promising performance of pre-trained vision-language models, e.g., CLIP, on versatile downstream tasks. In this paper, we focus on a particular setting of learning adaptive…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Chun-Mei Feng , Kai Yu , Yong Liu , Salman Khan , Wangmeng Zuo

A flexible recommendation and retrieval system requires music similarity in terms of multiple partial elements of musical pieces to allow users to select the element they want to focus on. A method for music similarity learning using…

Sound · Computer Science 2025-07-18 Yuka Hashizume , Li Li , Atsushi Miyashita , Tomoki Toda

Singing voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authenticity. Unlike Audio…

Sound · Computer Science 2023-09-07 Yuankun Xie , Jingjing Zhou , Xiaolin Lu , Zhenghao Jiang , Yuxin Yang , Haonan Cheng , Long Ye