English
Related papers

Related papers: The NES Music Database: A multi-instrumental datas…

200 papers

Cardiac auscultation, an integral tool in diagnosing cardiovascular diseases (CVDs), often relies on the subjective interpretation of clinicians, presenting a limitation in consistency and accuracy. Addressing this, we introduce the BUET…

Enhancing the ability of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) to interpret sheet music is a crucial step toward building AI musicians. However, current research lacks both evaluation benchmarks and…

Computation and Language · Computer Science 2025-09-29 Zhilin Wang , Zhe Yang , Yun Luo , Yafu Li , Xiaoye Qu , Ziqian Qiao , Haoran Zhang , Runzhe Zhan , Derek F. Wong , Jizhe Zhou , Yu Cheng

This paper presents a generative AI model for automated music composition with LSTM networks that takes a novel approach at encoding musical information which is based on movement in music rather than absolute pitch. Melodies are encoded as…

Sound · Computer Science 2021-08-25 Hooman Rafraf

Expressive music synthesis (EMS) for violin performance is a challenging task due to the disagreement among music performers in the interpretation of expressive musical terms (EMTs), scarcity of labeled recordings, and limited…

Sound · Computer Science 2024-06-27 Tzu-Yun Hung , Jui-Te Wu , Yu-Chia Kuo , Yo-Wei Hsiao , Ting-Wei Lin , Li Su

Developing digital sound synthesizers is crucial to the music industry as it provides a low-cost way to produce high-quality sounds with rich timbres. Existing traditional synthesizers often require substantial expertise to determine the…

Sound · Computer Science 2023-05-23 Zhen Ye , Wei Xue , Xu Tan , Qifeng Liu , Yike Guo

While recent years have seen remarkable progress in music generation models, research on their biases across countries, languages, cultures, and musical genres remains underexplored. This gap is compounded by the lack of datasets and…

Sound · Computer Science 2025-10-03 Ahmet Solak , Florian Grötschla , Luca A. Lanzendörfer , Roger Wattenhofer

One of the most significant challenges in Music Emotion Recognition (MER) comes from the fact that emotion labels can be heterogeneous across datasets with regard to the emotion representation, including categorical (e.g., happy, sad)…

Sound · Computer Science 2025-04-14 Jaeyong Kang , Dorien Herremans

Representation learning focused on disentangling the underlying factors of variation in given data has become an important area of research in machine learning. However, most of the studies in this area have relied on datasets from the…

Machine Learning · Computer Science 2020-07-31 Ashis Pati , Siddharth Gururani , Alexander Lerch

We consider the task of multimodal music mood prediction based on the audio signal and the lyrics of a track. We reproduce the implementation of traditional feature engineering based approaches and propose a new model based on deep…

Information Retrieval · Computer Science 2018-09-21 Rémi Delbouys , Romain Hennequin , Francesco Piccoli , Jimena Royo-Letelier , Manuel Moussallam

The automatic generation of medleys, i.e., musical pieces formed by different songs concatenated via smooth transitions, is not well studied in the current literature. To facilitate research on this topic, we make available a dataset called…

Sound · Computer Science 2020-08-26 Lukas Faber , Sandro Luck , Damian Pascual , Andreas Roth , Gino Brunner , Roger Wattenhofer

Separation of multiple singing voices into each voice is a rarely studied area in music source separation research. The absence of a benchmark dataset has hindered its progress. In this paper, we present an evaluation dataset and provide…

Sound · Computer Science 2023-05-05 Chang-Bin Jeon , Hyeongi Moon , Keunwoo Choi , Ben Sangbae Chon , Kyogu Lee

Music mixing involves combining individual tracks into a cohesive mixture, a task characterized by subjectivity where multiple valid solutions exist for the same input. Existing automatic mixing systems treat this task as a deterministic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-12 Eloi Moliner , Marco A. Martínez-Ramírez , Junghyun Koo , Wei-Hsiang Liao , Kin Wai Cheuk , Joan Serrà , Vesa Välimäki , Yuki Mitsufuji

In this work, we introduce the demonstration of symbolic music generation, focusing on providing short musical motifs that serve as the central theme of the narrative. For the generation, we adopt an autoregressive model which takes musical…

In this paper, we present a new dataset of music performance videos which can be used for training machine learning methods for multiple tasks such as audio-visual blind source separation and localization, cross-modal correspondences,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-10 Juan F. Montesinos , Olga Slizovskaia , Gloria Haro

Moonbeam is a transformer-based foundation model for symbolic music, pretrained on a large and diverse collection of MIDI data totaling 81.6K hours of music and 18 billion tokens. Moonbeam incorporates music-domain inductive biases by…

Sound · Computer Science 2025-05-22 Zixun Guo , Simon Dixon

We introduce MMIS, a novel dataset designed to advance MultiModal Interior Scene generation and recognition. MMIS consists of nearly 160,000 images. Each image within the dataset is accompanied by its corresponding textual description and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Hozaifa Kassab , Ahmed Mahmoud , Mohamed Bahaa , Ammar Mohamed , Ali Hamdi

High-quality datasets for learning-based modelling of polyphonic symbolic music remain less readily-accessible at scale than in other domains, such as language modelling or image classification. Deep learning algorithms show great potential…

Sound · Computer Science 2022-04-04 Omar Peracha

Deep learning models for music have advanced drastically in recent years, but how good are machine learning models at capturing emotion, and what challenges are researchers facing? In this paper, we provide a comprehensive overview of the…

Sound · Computer Science 2025-06-25 Jaeyong Kang , Dorien Herremans

Many machine learning algorithms represent input data with vector embeddings or discrete codes. When inputs exhibit compositional structure (e.g. objects built from parts or procedures from subroutines), it is natural to ask whether this…

Machine Learning · Computer Science 2019-04-09 Jacob Andreas

In the domain of algorithmic music composition, machine learning-driven systems eliminate the need for carefully hand-crafting rules for composition. In particular, the capability of recurrent neural networks to learn complex temporal…

Sound · Computer Science 2019-03-05 Harish Kumar , Balaraman Ravindran