中文
相关论文

相关论文: HSD: A hierarchical singing annotation dataset

200 篇论文

This paper aims to test whether a multi-modal approach for music emotion recognition (MER) performs better than a uni-modal one on high-level song features and lyrics. We use 11 song features retrieved from the Spotify API, combined lyrics…

声音 · 计算机科学 2023-02-28 Tibor Krols , Yana Nikolova , Ninell Oldenburg

Automatic music transcription converts audio recordings into symbolic representations, facilitating music analysis, retrieval, and generation. A musical note is characterized by pitch, onset, and offset in an audio domain, whereas it is…

声音 · 计算机科学 2025-02-19 Leekyung Kim , Sungwook Jeon , Wan Heo , Jonghun Park

In this work, we address the challenge of lyrics alignment, which involves aligning the lyrics and vocal components of songs. This problem requires the alignment of two distinct modalities, namely text and audio. To overcome this challenge,…

声音 · 计算机科学 2023-07-11 Minsung Kang , Soochul Park , Keunwoo Choi

Musical features and descriptors could be coarsely divided into three levels of complexity. The bottom level contains the basic building blocks of music, e.g., chords, beats and timbre. The middle level contains concepts that emerge from…

声音 · 计算机科学 2018-06-14 Anna Aljanaki , Mohammad Soleymani

The Complete Vocal Technique (CVT) is a school of singing developed in the past decades by Cathrin Sadolin et al.. CVT groups the use of the voice into so called vocal modes, namely Neutral, Curbing, Overdrive and Edge. Knowledge of the…

声音 · 计算机科学 2026-04-30 Reemt Hinrichs , Sonja Stephan , Alexander Lange , Jörn Ostermann

Conventional music structure analysis algorithms aim to divide a song into segments and to group them with abstract labels (e.g., 'A', 'B', and 'C'). However, explicitly identifying the function of each segment (e.g., 'verse' or 'chorus')…

音频与语音处理 · 电气工程与系统科学 2022-05-31 Ju-Chiang Wang , Yun-Ning Hung , Jordan B. L. Smith

This paper introduces a new large-scale music dataset, MusicNet, to serve as a source of supervision and evaluation of machine learning methods for music research. MusicNet consists of hundreds of freely-licensed classical music recordings…

机器学习 · 统计学 2017-04-07 John Thickstun , Zaid Harchaoui , Sham Kakade

Structure perception is a fundamental aspect of music cognition in humans. Historically, the hierarchical organization of music into structures served as a narrative device for conveying meaning, creating expectancy, and evoking emotions in…

声音 · 计算机科学 2023-03-28 Nicolas Lazzari , Andrea Poltronieri , Valentina Presutti

Creating a pop song melody according to pre-written lyrics is a typical practice for composers. A computational model of how lyrics are set as melodies is important for automatic composition systems, but an end-to-end lyric-to-melody model…

音频与语音处理 · 电气工程与系统科学 2023-01-05 Daiyu Zhang , Ju-Chiang Wang , Katerina Kosta , Jordan B. L. Smith , Shicen Zhou

In this paper, we present a new dataset of music performance videos which can be used for training machine learning methods for multiple tasks such as audio-visual blind source separation and localization, cross-modal correspondences,…

音频与语音处理 · 电气工程与系统科学 2020-08-10 Juan F. Montesinos , Olga Slizovskaia , Gloria Haro

Computational engine sound modeling is central to the automotive audio industry, particularly for active sound design, virtual prototyping, and emerging data-driven engine sound synthesis methods. These applications require large volumes of…

声音 · 计算机科学 2026-03-10 Robin Doerfler , Lonce Wyse

Automatic Singing Assessment and Singing Information Processing have evolved over the past three decades to support singing pedagogy, performance analysis, and vocal training. While the first approach objectively evaluates a singer's…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Arthur N. dos Santos , Bruno S. Masiero

One of the main limitations in the field of audio signal processing is the lack of large public datasets with audio representations and high-quality annotations due to restrictions of copyrighted commercial music. We present Melon Playlist…

The automatic generation of medleys, i.e., musical pieces formed by different songs concatenated via smooth transitions, is not well studied in the current literature. To facilitate research on this topic, we make available a dataset called…

声音 · 计算机科学 2020-08-26 Lukas Faber , Sandro Luck , Damian Pascual , Andreas Roth , Gino Brunner , Roger Wattenhofer

Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist due to the lack of a captioned MIDI dataset. This work aims to enable research that combines LLMs with symbolic…

音频与语音处理 · 电气工程与系统科学 2025-08-08 Jan Melechovsky , Abhinaba Roy , Dorien Herremans

Recent advances in deep learning have expanded possibilities to generate music, but generating a customizable full piece of music with consistent long-term structure remains a challenge. This paper introduces MusicFrameworks, a hierarchical…

声音 · 计算机科学 2021-09-03 Shuqi Dai , Zeyu Jin , Celso Gomes , Roger B. Dannenberg

Large-scale annotated datasets allow AI systems to learn from and build upon the knowledge of the crowd. Many crowdsourcing techniques have been developed for collecting image annotations. These techniques often implicitly rely on the fact…

人机交互 · 计算机科学 2016-10-07 Gunnar A. Sigurdsson , Olga Russakovsky , Ali Farhadi , Ivan Laptev , Abhinav Gupta

Artificial Intelligence Generated Content (AIGC) is currently a popular research area. Among its various branches, song generation has attracted growing interest. Despite the abundance of available songs, effective data preparation remains…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Wei Tan , Shun Lei , Huaicheng Zhang , Guangzheng Li , Yixuan Zhang , Hangting Chen , Jianwei Yu , Rongzhi Gu , Dong Yu

We propose a data cleansing method that utilizes a neural analysis and synthesis (NANSY++) framework to train an end-to-end neural diarization model (EEND) for singer diarization. Our proposed model converts song data with choral singing…

音频与语音处理 · 电气工程与系统科学 2024-06-25 Hokuto Munakata , Ryo Terashima , Yusuke Fujita

The HuggingFace Datasets Hub hosts thousands of datasets, offering exciting opportunities for language model training and evaluation. However, datasets for a specific task type often have different schemas, making harmonization challenging.…

计算与语言 · 计算机科学 2023-05-17 Damien Sileo