English
Related papers

Related papers: MOSA: Music Motion with Semantic Annotation Datase…

200 papers

Symbolic music is represented in two distinct forms: two-dimensional, visually intuitive score images, and one-dimensional, standardized text annotation sequences. While large language models have shown extraordinary potential in music,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Mingni Tang , Jiajia Li , Lu Yang , Zhiqiang Zhang , Jinghao Tian , Zuchao Li , Lefei Zhang , Ping Wang

We introduce a dataset for facilitating audio-visual analysis of music performances. The dataset comprises 44 simple multi-instrument classical music pieces assembled from coordinated but separately recorded performances of individual…

Multimedia · Computer Science 2018-08-09 Bochen Li , Xinzhao Liu , Karthik Dinesh , Zhiyao Duan , Gaurav Sharma

Music exists in various modalities, such as score images, symbolic scores, MIDI, and audio. Translations between each modality are established as core tasks of music information retrieval, such as automatic music transcription…

Sound · Computer Science 2026-04-08 Jongmin Jung , Dongmin Kim , Sihun Lee , Seola Cho , Hyungjoon Soh , Irmak Bukey , Chris Donahue , Dasaem Jeong

Multimodal sentiment analysis (MSA) draws increasing attention with the availability of multimodal data. The boost in performance of MSA models is mainly hindered by two problems. On the one hand, recent MSA works mostly focus on learning…

Machine Learning · Computer Science 2021-11-17 Ying Zeng , Sijie Mai , Haifeng Hu

Music is characterized by aspects related to different modalities, such as the audio signal, the lyrics, or the music video clips. This has motivated the development of multimodal datasets and methods for Music Information Retrieval (MIR)…

Multimedia · Computer Science 2025-09-19 Jonas Geiger , Marta Moscati , Shah Nawaz , Markus Schedl

This work present a music dataset named MusicTM-Dataset, which is utilized in improving the representation learning ability of different types of cross-modal retrieval (CMR). Little large music dataset including three modalities is…

Sound · Computer Science 2021-05-10 Donghuo Zeng , Yi Yu , Keizo Oyama

Multimodal learning has driven innovation across various industries, particularly in the field of music. By enabling more intuitive interaction experiences and enhancing immersion, it not only lowers the entry barriers to the music but also…

Multimedia · Computer Science 2026-02-24 Sifei Li , Mining Tan , Feier Shen , Minyan Luo , Zijiao Yin , Fan Tang , Weiming Dong , Changsheng Xu

Collecting large, aligned cross-modal datasets for music-flavor research is difficult because perceptual experiments are costly and small by design. We address this bottleneck through two complementary experiments. The first tests whether…

Sound · Computer Science 2026-04-14 Matteo Spanio , Valentina Frezzato , Antonio Rodà

The main challenges of Optical Music Recognition (OMR) come from the nature of written music, its complexity and the difficulty of finding an appropriate data representation. This paper provides a first look at DoReMi, an OMR dataset that…

Information Retrieval · Computer Science 2021-07-19 Elona Shatri , György Fazekas

Multi-modal learning has shown exceptional performance in various tasks, especially in medical applications, where it integrates diverse medical information for comprehensive diagnostic evidence. However, there still are several challenges…

Machine Learning · Computer Science 2024-11-19 Lin Fan , Yafei Ou , Cenyang Zheng , Pengyu Dai , Tamotsu Kamishima , Masayuki Ikebe , Kenji Suzuki , Xun Gong

Multimodal Sentiment Analysis (MSA) aims to predict sentiment from language, acoustic, and visual data in videos. However, imbalanced unimodal performance often leads to suboptimal fused representations. Existing approaches typically adopt…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Dingkang Yang , Mingcheng Li , Xuecheng Wu , Zhaoyu Chen , Kaixun Jiang , Keliang Liu , Peng Zhai , Lihua Zhang

This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Zitang Zhou , Ke Mei , Yu Lu , Tianyi Wang , Fengyun Rao

While piano music has become a significant area of study in Music Information Retrieval (MIR), there is a notable lack of datasets for piano solo music with text labels. To address this gap, we present PIAST (PIano dataset with Audio,…

Sound · Computer Science 2024-11-08 Hayeon Bang , Eunjin Choi , Megan Finch , Seungheon Doh , Seolhee Lee , Gyeong-Hoon Lee , Juhan Nam

While multi-modal learning has advanced significantly, current approaches often treat modalities separately, creating inconsistencies in representation and reasoning. We introduce MANTA (Multi-modal Abstraction and Normalization via Textual…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Ziqi Zhong , Daniel Tang

Optical Music Recognition (OMR) has long been without an adequate dataset and ground truth for evaluating OMR systems, which has been a major problem for establishing a state of the art in the field. Furthermore, machine learning methods…

Computer Vision and Pattern Recognition · Computer Science 2017-03-16 Jan Hajič , Pavel Pecina

There has been a rapid growth of digitally available music data, including audio recordings, digitized images of sheet music, album covers and liner notes, and video clips. This huge amount of data calls for retrieval strategies that allow…

Information Retrieval · Computer Science 2019-02-13 Meinard Müller , Andreas Arzt , Stefan Balke , Matthias Dorfer , Gerhard Widmer

Research on large language models has advanced significantly across text, speech, images, and videos. However, multi-modal music understanding and generation remain underexplored due to the lack of well-annotated datasets. To address this,…

Sound · Computer Science 2024-12-10 Shansong Liu , Atin Sakkeer Hussain , Qilong Wu , Chenshuo Sun , Ying Shan

Code-switching, the alternation between two or more languages within communication, poses great challenges for Automatic Speech Recognition (ASR) systems. Existing models and datasets are limited in their ability to effectively handle these…

Sound · Computer Science 2025-11-14 Yupei Li , Zifan Wei , Heng Yu , Jiahao Xue , Huichi Zhou , Björn W. Schuller

Storytelling is multi-modal in the real world. When one tells a story, one may use all of the visualizations and sounds along with the story itself. However, prior studies on storytelling datasets and tasks have paid little attention to…

Multimedia · Computer Science 2023-10-31 Jaeyeon Bae , Seokhoon Jeong , Seokun Kang , Namgi Han , Jae-Yon Lee , Hyounghun Kim , Taehwan Kim

Music representation learning is central to music information retrieval and generation. While recent advances in multimodal learning have improved alignment between text and audio for tasks such as cross-modal music retrieval, text-to-music…

‹ Prev 1 2 3 10 Next ›