English
Related papers

Related papers: NOTA: Multimodal Music Notation Understanding for …

200 papers

We propose a framework for audio-to-score alignment on piano performance that employs automatic music transcription (AMT) using neural networks. Even though the AMT result may contain some errors, the note prediction output can be regarded…

Sound · Computer Science 2017-11-15 Taegyun Kwon , Dasaem Jeong , Juhan Nam

This work addresses the problem of matching short excerpts of audio with their respective counterparts in sheet music images. We show how to employ neural network-based cross-modality embedding spaces for solving the following two sheet…

Information Retrieval · Computer Science 2017-08-01 Matthias Dorfer , Andreas Arzt , Gerhard Widmer

Could we automatically derive the score of a piano accompaniment based on the audio of a pop song? This is the audio-to-symbolic arrangement problem we tackle in this paper. A good arrangement model should not only consider the audio…

Sound · Computer Science 2022-02-23 Ziyu Wang , Dejing Xu , Gus Xia , Ying Shan

Charts play a vital role in data visualization, understanding data patterns, and informed decision-making. However, their unique combination of graphical elements (e.g., bars, lines) and textual components (e.g., labels, legends) poses…

Computer Vision and Pattern Recognition · Computer Science 2024-02-16 Fanqing Meng , Wenqi Shao , Quanfeng Lu , Peng Gao , Kaipeng Zhang , Yu Qiao , Ping Luo

We propose a method to recommend music for an input video while allowing a user to guide music selection with free-form natural language. A key challenge of this problem setting is that existing music video datasets provide the needed…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Daniel McKee , Justin Salamon , Josef Sivic , Bryan Russell

Existing music captioning methods are limited to generating concise global descriptions of short music clips, which fail to capture fine-grained musical characteristics and time-aware musical changes. To address these limitations, we…

Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generating music that…

Sound · Computer Science 2024-09-13 Tanisha Hisariya , Huan Zhang , Jinhua Liang

Properly annotated multimedia content is crucial for supporting advances in many Information Retrieval applications. It enables, for instance, the development of automatic tools for the annotation of large and diverse multimedia…

Information Retrieval · Computer Science 2018-11-28 Xavier Favory , Eduardo Fonseca , Frederic Font , Xavier Serra

A representation technique that allows encoding music in a way that contains musical meaning would improve the results of any model trained for computer music tasks like generation of melodies and harmonies of better quality. The field of…

Computation and Language · Computer Science 2020-05-20 Sebastian Garcia-Valencia

We present PandaGPT, an approach to emPower large lANguage moDels with visual and Auditory instruction-following capabilities. Our pilot experiments show that PandaGPT can perform complex tasks such as detailed image description generation,…

Computation and Language · Computer Science 2023-05-29 Yixuan Su , Tian Lan , Huayang Li , Jialu Xu , Yan Wang , Deng Cai

Sign language recognition is a challenging and often underestimated problem comprising multi-modal articulators (handshape, orientation, movement, upper body and face) that integrate asynchronously on multiple streams. Learning powerful…

Computer Vision and Pattern Recognition · Computer Science 2019-11-22 Hamid Reza Vaezi Joze , Oscar Koller

This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Baoyao Yang , Wanyun Li , Dixin Chen , Junxiang Chen , Wenbin Yao , Haifeng Lin

Optical music recognition (OMR) aims to convert music notation into digital formats. One approach to tackle OMR is through a multi-stage pipeline, where the system first detects visual music notation elements in the image (object detection)…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Guang Yang , Muru Zhang , Lin Qiu , Yanming Wan , Noah A. Smith

Recent years have seen many audio-domain text-to-music generation models that rely on large amounts of text-audio pairs for training. However, symbolic-domain controllable music generation has lagged behind partly due to the lack of a…

Sound · Computer Science 2025-06-17 Weihan Xu , Julian McAuley , Taylor Berg-Kirkpatrick , Shlomo Dubnov , Hao-Wen Dong

Pre-trained large language models have recently achieved ground-breaking performance in a wide variety of language understanding tasks. However, the same model can not be applied to multimodal behavior understanding tasks (e.g., video…

Computation and Language · Computer Science 2023-03-30 Md Kamrul Hasan , Md Saiful Islam , Sangwu Lee , Wasifur Rahman , Iftekhar Naim , Mohammed Ibrahim Khan , Ehsan Hoque

Current approaches to music emotion annotation remain heavily reliant on manual labelling, a process that imposes significant resource and labour burdens, severely limiting the scale of available annotated data. This study examines the…

Sound · Computer Science 2025-08-19 Meng Yang , Jon McCormack , Maria Teresa Llano , Wanchao Su

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Jielin Qiu , Jiacheng Zhu , William Han , Aditesh Kumar , Karthik Mittal , Claire Jin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Ding Zhao , Bo Li , Lijuan Wang

Multi-modal large language models have demonstrated impressive performance across various tasks in different modalities. However, existing multi-modal models primarily emphasize capturing global information within each modality while…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Zhaowei Li , Qi Xu , Dong Zhang , Hang Song , Yiqing Cai , Qi Qi , Ran Zhou , Junting Pan , Zefeng Li , Van Tu Vu , Zhida Huang , Tao Wang

Many music AI models learn a map between music content and human-defined labels. However, many annotations, such as chords, can be naturally expressed within the music modality itself, e.g., as sequences of symbolic notes. This observation…

Sound · Computer Science 2025-09-30 Junyan Jiang , Daniel Chin , Liwei Lin , Xuanjie Liu , Gus Xia

Scaling semantic parsing models for task-oriented dialog systems to new languages is often expensive and time-consuming due to the lack of available datasets. Available datasets suffer from several shortcomings: a) they contain few…

Computation and Language · Computer Science 2021-01-28 Haoran Li , Abhinav Arora , Shuohui Chen , Anchit Gupta , Sonal Gupta , Yashar Mehdad
‹ Prev 1 4 5 6 7 8 10 Next ›