中文
相关论文

相关论文: Melody transcription via generative pre-training

200 篇论文

Extraction of predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled audio…

音频与语音处理 · 电气工程与系统科学 2024-02-13 Kavya Ranjan Saxena , Vipul Arora

In the realm of music information retrieval, similarity-based retrieval and auto-tagging serve as essential components. Given the limitations and non-scalability of human supervision signals, it becomes crucial for models to learn from…

Music-to-dance translation is a brand-new and powerful feature in recent role-playing games. Players can now let their characters dance along with specified music clips and even generate fan-made dance videos. Previous works of this topic…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Yinglin Duan , Tianyang Shi , Zhengxia Zou , Jia Qin , Yifei Zhao , Yi Yuan , Jie Hou , Xiang Wen , Changjie Fan

We propose MoodNet - A Deep Convolutional Neural Network based architecture to effectively predict the emotion associated with a piece of music given its audio and lyrical content.We evaluate different architectures consisting of varying…

音频与语音处理 · 电气工程与系统科学 2018-11-15 Aniruddha Bhattacharya , K. V. Kadambari

We examine the problem of learning a probabilistic model for melody directly from musical sequences belonging to the same genre. This is a challenging task as one needs to capture not only the rich temporal structure evident in music, but…

机器学习 · 计算机科学 2012-07-03 Athina Spiliopoulou , Amos Storkey

Symbolic Music Generation relies on the contextual representation capabilities of the generative model, where the most prevalent approach is the Transformer-based model. The learning of musical context is also related to the structural…

声音 · 计算机科学 2022-07-12 Guowei Wu , Shipei Liu , Xiaoya Fan

Definitive embeddings remain a fundamental challenge of computational musicology for symbolic music in deep learning today. Analogous to natural language, music can be modeled as a sequence of tokens. This motivates the majority of existing…

声音 · 计算机科学 2020-10-19 Hongru Liang , Wenqiang Lei , Paul Yaozhu Chan , Zhenglu Yang , Maosong Sun , Tat-Seng Chua

In this work, we provide a broad comparative analysis of strategies for pre-training audio understanding models for several tasks in the music domain, including labelling of genre, era, origin, mood, instrumentation, key, pitch, vocal…

The advancement of machine learning in audio analysis has opened new possibilities for technology-enhanced music education. This paper introduces a framework for automatic singing mistake detection in the context of music pedagogy,…

音频与语音处理 · 电气工程与系统科学 2026-02-09 Sumit Kumar , Suraj Jaiswal , Parampreet Singh , Vipul Arora

Extraction of the predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled…

音频与语音处理 · 电气工程与系统科学 2023-04-07 Kavya Ranjan Saxena , Vipul Arora

This paper presents the Jazz Transformer, a generative model that utilizes a neural sequence model called the Transformer-XL for modeling lead sheets of Jazz music. Moreover, the model endeavors to incorporate structural events present in…

声音 · 计算机科学 2020-08-05 Shih-Lun Wu , Yi-Hsuan Yang

Symbolic music generation aims to create musical notes, which can help users compose music, such as generating target instrument tracks based on provided source tracks. In practical scenarios where there's a predefined ensemble of tracks…

声音 · 计算机科学 2023-10-02 Ang Lv , Xu Tan , Peiling Lu , Wei Ye , Shikun Zhang , Jiang Bian , Rui Yan

In this paper, we propose an efficient and reproducible deep learning model for musical onset detection (MOD). We first review the state-of-the-art deep learning models for MOD, and identify their shortcomings and challenges: (i) the lack…

声音 · 计算机科学 2018-06-20 Rong Gong , Xavier Serra

Text-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this process demands extensive, high-quality data, which are often…

声音 · 计算机科学 2025-06-18 Chang Li , Ruoyu Wang , Lijuan Liu , Jun Du , Yixuan Sun , Zilu Guo , Zhenrong Zhang , Yuan Jiang , Jianqing Gao , Feng Ma

The quantity of processed data is crucial for advancing the field of singing voice synthesis. While there are tools available for lyric or note transcription tasks, they all need pre-processed data which is relatively time-consuming (e.g.,…

声音 · 计算机科学 2024-10-11 Siwei Wu , Jinzheng He , Ruibin Yuan , Haojie Wei , Xipin Wei , Chenghua Lin , Jin Xu , Junyang Lin

Deep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and…

声音 · 计算机科学 2021-02-08 Ho-Hsiang Wu , Chieh-Chi Kao , Qingming Tang , Ming Sun , Brian McFee , Juan Pablo Bello , Chao Wang

Modelling musical structure is vital yet challenging for artificial intelligence systems that generate symbolic music compositions. This literature review dissects the evolution of techniques for incorporating coherent structure, from…

声音 · 计算机科学 2024-03-14 Keshav Bhandari , Simon Colton

Symbolic Music Emotion Recognition(SMER) is to predict music emotion from symbolic data, such as MIDI and MusicXML. Previous work mainly focused on learning better representation via (mask) language model pre-training but ignored the…

声音 · 计算机科学 2022-01-19 Jibao Qiu , C. L. Philip Chen , Tong Zhang

Pitch and meter are two fundamental music features for symbolic music generation tasks, where researchers usually choose different encoding methods depending on specific goals. However, the advantages and drawbacks of different encoding…

声音 · 计算机科学 2023-02-01 Yuqiang Li , Shengchen Li , George Fazekas

We show that pre-training a Transformer on music before language significantly accelerates language acquisition. Using piano performances (MAESTRO dataset), a developmental pipeline -- music $\to$ poetry $\to$ prose -- yields a $17.5\%$…

计算与语言 · 计算机科学 2026-04-24 Yoshinori Nomura