中文
相关论文

相关论文: End-to-end learning for music audio tagging at sca…

200 篇论文

Machine learning approaches to auditory object recognition are traditionally based on engineered features such as those derived from the spectrum or cepstrum. More recently, end-to-end classification systems in image and auditory…

Word embedding has become an essential means for text-based information retrieval. Typically, word embeddings are learned from large quantities of general and unstructured text data. However, in the domain of music, the word embedding may…

声音 · 计算机科学 2024-04-24 SeungHeon Doh , Jongpil Lee , Dasaem Jeong , Juhan Nam

Recently, we proposed a self-attention based music tagging model. Different from most of the conventional deep architectures in music information retrieval, which use stacked 3x3 filters by treating music spectrograms as images, the…

声音 · 计算机科学 2019-11-12 Minz Won , Sanghyuk Chun , Xavier Serra

Large language models perform strongly on general tasks but remain constrained in specialized settings such as music, particularly in the music-entertainment domain, where corpus scale, purity, and the match between data and training…

计算与语言 · 计算机科学 2025-11-19 Kai Tian , Yirong Mao , Wendong Bi , Hanjie Wang , Que Wenhui

We propose MoodNet - A Deep Convolutional Neural Network based architecture to effectively predict the emotion associated with a piece of music given its audio and lyrical content.We evaluate different architectures consisting of varying…

音频与语音处理 · 电气工程与系统科学 2018-11-15 Aniruddha Bhattacharya , K. V. Kadambari

Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, while academic research remains largely non-reproducible due to the lack of publicly available training data, hindering fair comparison and…

This paper introduces effective design choices for text-to-music retrieval systems. An ideal text-based retrieval system would support various input queries such as pre-defined tags, unseen tags, and sentence-level descriptions. In reality,…

信息检索 · 计算机科学 2022-11-29 SeungHeon Doh , Minz Won , Keunwoo Choi , Juhan Nam

Learning music representations that are general-purpose offers the flexibility to finetune several downstream tasks using smaller datasets. The wav2vec 2.0 speech representation model showed promising results in many downstream speech…

音频与语音处理 · 电气工程与系统科学 2022-10-28 Alessandro Ragano , Emmanouil Benetos , Andrew Hines

We showcase an unsupervised method that repurposes deep models trained for music generation and music tagging for audio source separation, without any retraining. An audio generation model is conditioned on an input mixture, producing a…

声音 · 计算机科学 2021-10-26 Ethan Manilow , Patrick O'Reilly , Prem Seetharaman , Bryan Pardo

Conventional music structure analysis algorithms aim to divide a song into segments and to group them with abstract labels (e.g., 'A', 'B', and 'C'). However, explicitly identifying the function of each segment (e.g., 'verse' or 'chorus')…

音频与语音处理 · 电气工程与系统科学 2022-05-31 Ju-Chiang Wang , Yun-Ning Hung , Jordan B. L. Smith

Many practices have been presented in music generation recently. While stylistic music generation using deep learning techniques has became the main stream, these models still struggle to generate music with high musicality, different…

声音 · 计算机科学 2021-05-12 Shuqi Dai , Xichu Ma , Ye Wang , Roger B. Dannenberg

In this work, we address music representation learning using convolution-free transformers. We build on top of existing spectrogram-based audio transformers such as AST and train our models on a supervised task using patchout training…

声音 · 计算机科学 2023-09-29 Pablo Alonso-Jiménez , Xavier Serra , Dmitry Bogdanov

Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large language models have shown extraordinary potential in music,…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Mingni Tang , Jiajia Li , Lu Yang , Zhiqiang Zhang , Jinghao Tian , Zuchao Li , Lefei Zhang , Ping Wang

We propose a knowledge-driven, model-based approach to segmenting audio into single-category and mixed-category chunks with applications to source separation. "Knowledge" here denotes information associated with the data, such as music…

音频与语音处理 · 电气工程与系统科学 2026-02-26 Chun-wei Ho , Sabato Marco Siniscalchi , Kai Li , Chin-Hui Lee

We have seen remarkable success in representation learning and language models (LMs) using deep neural networks. Many studies aim to build the underlying connections among different modalities via the alignment and mappings at the token or…

声音 · 计算机科学 2025-03-04 Daniel Chin , Gus Xia

In this paper, we propose a framework for environmental sound classification in a low-data context (less than 100 labeled examples per class). We show that using pre-trained image classification models along with the usage of data…

声音 · 计算机科学 2019-09-30 Sainath Adapa

Generative models in vision have seen rapid progress due to algorithmic improvements and the availability of high-quality image datasets. In this paper, we offer contributions in both these areas to enable similar progress in audio…

机器学习 · 计算机科学 2017-04-06 Jesse Engel , Cinjon Resnick , Adam Roberts , Sander Dieleman , Douglas Eck , Karen Simonyan , Mohammad Norouzi

Recent progress in deep learning for audio synthesis opens the way to models that directly produce the waveform, shifting away from the traditional paradigm of relying on vocoders or MIDI synthesizers for speech or music generation. Despite…

声音 · 计算机科学 2018-10-24 Alexandre Défossez , Neil Zeghidour , Nicolas Usunier , Léon Bottou , Francis Bach

We apply deep learning methods, specifically long short-term memory (LSTM) networks, to music transcription modelling and composition. We build and train LSTM networks using approximately 23,000 music transcriptions expressed with a…

声音 · 计算机科学 2016-05-02 Bob L. Sturm , João Felipe Santos , Oded Ben-Tal , Iryna Korshunova

With the growing amount of musical data available, automatic instrument recognition, one of the essential problems in Music Information Retrieval (MIR), is drawing more and more attention. While automatic recognition of single instruments…

声音 · 计算机科学 2023-06-16 Lifan Zhong , Erica Cooper , Junichi Yamagishi , Nobuaki Minematsu