English
Related papers

Related papers: Two Web Toolkits for Multimodal Piano Performance …

200 papers

We present Piano Genie, an intelligent controller which allows non-musicians to improvise on the piano. With Piano Genie, a user performs on a simple interface with eight buttons, and their performance is decoded into the space of plausible…

Machine Learning · Computer Science 2019-03-25 Chris Donahue , Ian Simon , Sander Dieleman

We present a demonstration of a web-based system called M2LADS ("System for Generating Multimodal Learning Analytics Dashboards"), designed to integrate, synchronize, visualize, and analyze multimodal data recorded during computer-based…

Human-Computer Interaction · Computer Science 2025-03-17 Alvaro Becerra , Roberto Daza , Ruth Cobos , Aythami Morales , Julian Fierrez

Speech production is a complex process spanning neural planning, motor control, muscle activation, and articulatory kinematics. While the acoustic speech signal is the most accessible product of the speech production act, it does not…

Feature alignment serves as the primary mechanism for fusing multimodal data. We put forth a feature alignment approach that achieves full integration of multimodal information. This is accomplished via an alternating process of shifting…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Jiahao Qin

This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio…

Sound · Computer Science 2025-07-30 Ohad Cohen , Gershon Hazan , Sharon Gannot

This paper addresses the problem of sheet-image-based on-line audio-to-score alignment also known as score following. Drawing inspiration from object detection, a conditional neural network architecture is proposed that directly predicts…

Sound · Computer Science 2021-05-11 Florian Henkel , Gerhard Widmer

We present PKSpell: a data-driven approach for the joint estimation of pitch spelling and key signatures from MIDI files. Both elements are fundamental for the production of a full-fledged musical score and facilitate many MIR tasks such as…

Sound · Computer Science 2021-07-30 Francesco Foscarin , Nicolas Audebert , Raphaël Fournier-S'Niehotta

In the realm of music AI, arranging rich and structured multi-track accompaniments from a simple lead sheet presents significant challenges. Such challenges include maintaining track cohesion, ensuring long-term coherence, and optimizing…

Sound · Computer Science 2024-11-26 Jingwei Zhao , Gus Xia , Ziyu Wang , Ye Wang

One of the factors that have hindered progress in the areas of sign language recognition, translation, and production is the absence of large annotated datasets. Towards this end, we introduce How2Sign, a multimodal and multiview continuous…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Amanda Duarte , Shruti Palaskar , Lucas Ventura , Deepti Ghadiyaram , Kenneth DeHaan , Florian Metze , Jordi Torres , Xavier Giro-i-Nieto

With the development of multimedia systems, multimodal recommendations are playing an essential role, as they can leverage rich contexts beyond interactions. Existing methods mainly regard multimodal information as an auxiliary, using them…

Information Retrieval · Computer Science 2024-08-02 Yifan Liu , Kangning Zhang , Xiangyuan Ren , Yanhua Huang , Jiarui Jin , Yingjie Qin , Ruilong Su , Ruiwen Xu , Yong Yu , Weinan Zhang

We revisit the problems of pitch spelling and tonality guessing with a new algorithm for their joint estimation from a MIDI file including information about the measure boundaries. Our algorithm does not only identify a global key but also…

Sound · Computer Science 2024-02-19 Augustin Bouquillard , Florent Jacquemard

Emotion estimation in music listening is confronting challenges to capture the emotion variation of listeners. Recent years have witnessed attempts to exploit multimodality fusing information from musical contents and physiological signals…

Artificial Intelligence · Computer Science 2016-12-01 Nattapong Thammasan , Ken-ichi Fukui , Masayuki Numao

Instrument playing is among the most common scenes in music-related videos, which represent nowadays one of the largest sources of online videos. In order to understand the instrument-playing scenes in the videos, it is important to know…

Multimedia · Computer Science 2018-05-08 Jen-Yu Liu , Yi-Hsuan Yang , Shyh-Kang Jeng

In the last years, scientific and industrial research has experienced a growing interest in acquiring large annotated data sets to train artificial intelligence algorithms for tackling problems in different domains. In this context, we have…

Human-Computer Interaction · Computer Science 2021-03-09 Silvio Barra , Salvatore M. Carta , Alessandro Giuliani , Alessia Pisu , Alessandro Sebastian Podda , DanieleRiboni

Musical dynamics form a core part of expressive singing voice performances. However, automatic analysis of musical dynamics for singing voice has received limited attention partly due to the scarcity of suitable datasets and a lack of clear…

Sound · Computer Science 2024-10-29 Jyoti Narang , Nazif Can Tamer , Viviana De La Vega , Xavier Serra

Detecting and interpreting operator actions, engagement, and object interactions in dynamic industrial workflows remains a significant challenge in human-robot collaboration research, especially within complex, real-world environments.…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Naval Kishore Mehta , Arvind , Himanshu Kumar , Abeer Banerjee , Sumeet Saurav , Sanjay Singh

the substantial increase in the number of online human-human conversations and the usefulness of multimodal transcripts, there is a rising need for automated multimodal transcription systems to help us better understand the conversations.…

Human-Computer Interaction · Computer Science 2022-01-03 Joshua Y. Kim , Kalina Yacef

Making a slight mistake during live music performance can easily be spotted by an astute listener, even if the performance is an improvisation or an unfamiliar piece. An example might be a highly dissonant chord played by mistake in a…

Sound · Computer Science 2020-12-01 Georgi Marinov

In this paper, we introduce Jointist, an instrument-aware multi-instrument framework that is capable of transcribing, recognizing, and separating multiple musical instruments from an audio clip. Jointist consists of an instrument…

Multi-instrument recognition is the task of predicting the presence or absence of different instruments within an audio clip. A considerable challenge in applying deep learning to multi-instrument recognition is the scarcity of labeled…

Sound · Computer Science 2020-01-27 Amir Kenarsari Anhari