中文
相关论文

相关论文: Multi-Modal Chorus Recognition for Improving Song …

200 篇论文

We introduce a novel low level feature for identifying cover songs which quantifies the relative changes in the smoothed frequency spectrum of a song. Our key insight is that a sliding window representation of a chunk of audio can be viewed…

声音 · 计算机科学 2015-07-21 Christopher J. Tralie , Paul Bendich

Little research focuses on cross-modal correlation learning where temporal structures of different data modalities such as audio and lyrics are taken into account. Stemming from the characteristic of temporal structures of music in nature,…

信息检索 · 计算机科学 2017-11-30 Yi Yu , Suhua Tang , Francisco Raposo , Lei Chen

Songwriting is often driven by multimodal inspirations, such as imagery, narratives, or existing music, yet songwriters remain unsupported by current music AI systems in incorporating these multimodal inputs into their creative processes.…

人机交互 · 计算机科学 2025-02-17 Yewon Kim , Sung-Ju Lee , Chris Donahue

In this work, we study the association between song lyrics and mood through a data-driven analysis. Our data set consists of nearly one million songs, with song-mood associations derived from user playlists on the Spotify streaming…

多媒体 · 计算机科学 2022-07-13 Shahrzad Naseri , Sravana Reddy , Joana Correia , Jussi Karlgren , Rosie Jones

Music source separation is a core task in music information retrieval which has seen a dramatic improvement in the past years. Nevertheless, most of the existing systems focus exclusively on the problem of source separation itself and…

音频与语音处理 · 电气工程与系统科学 2020-08-04 Yun-Ning Hung , Alexander Lerch

Previous approaches in singer identification have used one of monophonic vocal tracks or mixed tracks containing multiple instruments, leaving a semantic gap between these two domains of audio. In this paper, we present a system to learn a…

声音 · 计算机科学 2019-06-27 Kyungyun Lee , Juhan Nam

In this paper, we propose an efficient and reproducible deep learning model for musical onset detection (MOD). We first review the state-of-the-art deep learning models for MOD, and identify their shortcomings and challenges: (i) the lack…

声音 · 计算机科学 2018-06-20 Rong Gong , Xavier Serra

The central idea of this paper is to gain a deeper understanding of song lyrics computationally. We focus on two aspects: style and biases of song lyrics. All prior works to understand these two aspects are limited to manual analysis of a…

信息检索 · 计算机科学 2019-07-19 Manash Pratim Barman , Amit Awekar , Sambhav Kothari

This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio…

声音 · 计算机科学 2025-07-30 Ohad Cohen , Gershon Hazan , Sharon Gannot

This paper introduces a project of advanced system of music retrieval from the Internet. The system uses combination of text search (by author, title and other information about the music file included in id3 tag description or similar for…

信息检索 · 计算机科学 2013-09-18 M. Brzeziński-Spiczak , K. Dobosz , M. Lis , M. Pintal

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility,…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Jielin Qiu , Jiacheng Zhu , William Han , Aditesh Kumar , Karthik Mittal , Claire Jin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Ding Zhao , Bo Li , Lijuan Wang

We introduce a dataset for facilitating audio-visual analysis of music performances. The dataset comprises 44 simple multi-instrument classical music pieces assembled from coordinated but separately recorded performances of individual…

多媒体 · 计算机科学 2018-08-09 Bochen Li , Xinzhao Liu , Karthik Dinesh , Zhiyao Duan , Gaurav Sharma

Commonly music has an obvious hierarchical structure, especially for the singing parts which usually act as the main melody in pop songs. However, most of the current singing annotation datasets only record symbolic information of music…

声音 · 计算机科学 2022-10-03 Xiao Fu , Xin Yuan , Jinglu Hu

Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model's ability to analyze and interpret various musical elements. These improvements primarily focused on…

声音 · 计算机科学 2025-09-24 Zhuoyuan Mao , Mengjie Zhao , Qiyu Wu , Hiromi Wakaki , Yuki Mitsufuji

The quantity of processed data is crucial for advancing the field of singing voice synthesis. While there are tools available for lyric or note transcription tasks, they all need pre-processed data which is relatively time-consuming (e.g.,…

声音 · 计算机科学 2024-10-11 Siwei Wu , Jinzheng He , Ruibin Yuan , Haojie Wei , Xipin Wei , Chenghua Lin , Jin Xu , Junyang Lin

Chord recognition serves as a critical task in music information retrieval due to the abstract and descriptive nature of chords in music analysis. While audio chord recognition systems have achieved significant accuracy for small…

Sentiment prediction of contemporary music can have a wide-range of applications in modern society, for instance, selecting music for public institutions such as hospitals or restaurants to potentially improve the emotional well-being of…

机器学习 · 计算机科学 2016-11-02 Sebastian Raschka

This paper explores a specific sub-task of cross-modal music retrieval. We consider the delicate task of retrieving a performance or rendition of a musical piece based on a description of its style, expressive character, or emotion from a…

声音 · 计算机科学 2024-01-29 Shreyan Chowdhury , Gerhard Widmer

Many music AI models learn a map between music content and human-defined labels. However, many annotations, such as chords, can be naturally expressed within the music modality itself, e.g., as sequences of symbolic notes. This observation…

声音 · 计算机科学 2025-09-30 Junyan Jiang , Daniel Chin , Liwei Lin , Xuanjie Liu , Gus Xia

Music similarity retrieval is fundamental for managing and exploring relevant content from large collections in streaming platforms. This paper presents a novel cross-modal contrastive learning framework that leverages the open-ended nature…

声音 · 计算机科学 2025-05-26 Tristan Tsoi , Jiajun Deng , Yaolong Ju , Benno Weck , Holger Kirchhoff , Simon Lui