English
Related papers

Related papers: Decoding and Visualising Intended Emotion in an Ex…

200 papers

We present Affect2MM, a learning method for time-series emotion prediction for multimedia content. Our goal is to automatically capture the varying emotions depicted by characters in real-life human-centric situations and behaviors. We use…

Computer Vision and Pattern Recognition · Computer Science 2021-03-12 Trisha Mittal , Puneet Mathur , Aniket Bera , Dinesh Manocha

The primary objective is to teach a machine about human emotions, which has become an essential requirement in the field of social intelligence, also expedites the progress of human-machine interactions. The ability of a machine to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-23 Sai Nikhil Chennoor , B. R. K. Madhur , Moujiz Ali , T. Kishore Kumar

Expressive performance rendering (EPR) and automatic piano transcription (APT) are fundamental yet inverse tasks in music information retrieval: EPR generates expressive performances from symbolic scores, while APT recovers scores from…

Sound · Computer Science 2025-09-30 Wei Zeng , Junchuan Zhao , Ye Wang

Many music theoretical constructs (such as scale types, modes, cadences, and chord types) are defined in terms of pitch intervals---relative distances between pitches. Therefore, when computer models are employed in music tasks, it can be…

Sound · Computer Science 2019-02-05 Stefan Lattner , Maarten Grachten , Gerhard Widmer

Spatio-temporal feature encoding is essential for encoding facial expression dynamics in video sequences. At test time, most spatio-temporal encoding methods assume that a temporally segmented sequence is fed to a learned model, which could…

Computer Vision and Pattern Recognition · Computer Science 2017-11-30 Wissam J. Baddar , Yong Man Ro

The way that humans encode their emotion into speech signals is complex. For instance, an angry man may increase his pitch and speaking rate, and use impolite words. In this paper, we present a preliminary study on various emotional factors…

Sound · Computer Science 2021-11-25 Haoran Sun , Lantian Li , Thomas Fang Zheng , Dong Wang

Due to the complex nature of human emotions and the diversity of emotion representation methods in humans, emotion recognition is a challenging field. In this research, three input modalities, namely text, audio (speech), and video, are…

Artificial Intelligence · Computer Science 2024-02-13 Minoo Shayaninasab , Bagher Babaali

Our study investigates an approach for understanding musical performances through the lens of audio encoding models, focusing on the domain of solo Western classical piano music. Compared to composition-level attribute understanding such as…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-22 Huan Zhang , Jinhua Liang , Simon Dixon

Our study has investigated the effect of music on the experience of viewing art, investigating the factors which create a sense of connectivity between the two forms. We worked with 138 participants, and included multiple choice and…

Human-Computer Interaction · Computer Science 2025-01-10 Paul Warren , Paul Mulholland , Naomi Barker

Interaction with the world requires an organism to transform sensory signals into representations in which behaviorally meaningful properties of the environment are made explicit. These representations are derived through cascades of…

Neurons and Cognition · Quantitative Biology 2017-10-17 Wiktor Młynarski , Josh H. McDermott

Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist due to the lack of a captioned MIDI dataset. This work aims to enable research that combines LLMs with symbolic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-08 Jan Melechovsky , Abhinaba Roy , Dorien Herremans

We present a new large-scale emotion-labeled symbolic music dataset consisting of 12k MIDI songs. To create this dataset, we first trained emotion classification models on the GoEmotions dataset, achieving state-of-the-art results with a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-28 Serkan Sulun , Pedro Oliveira , Paula Viana

Evaluation for continuous piano pedal depth estimation tasks remains incomplete when relying only on conventional frame-level metrics, which overlook musically important features such as direction-change boundaries and pedal curve contours.…

Information Retrieval · Computer Science 2026-02-04 Hanwen Zhang , Kun Fang , Ziyu Wang , Ichiro Fujinaga

Emotion and intent recognition from speech is essential and has been widely investigated in human-computer interaction. The rapid development of social media platforms, chatbots, and other technologies has led to a large volume of speech…

Sound · Computer Science 2025-07-11 Zhao Ren , Rathi Adarshi Rammohan , Kevin Scheck , Sheng Li , Tanja Schultz

Research on style transfer and domain translation has clearly demonstrated the ability of deep learning-based algorithms to manipulate images in terms of artistic style. More recently, several attempts have been made to extend such…

Sound · Computer Science 2021-06-11 Ondřej Cífka , Umut Şimşekli , Gaël Richard

Technological advancement and its omnipresent connection have pushed humans past the boundaries and limitations of a computer screen, physical state, or geographical location. It has provided a depth of avenues that facilitate…

Multimedia · Computer Science 2023-11-21 Dayo Samuel Banjo , Connice Trimmingham , Niloofar Yousefi , Nitin Agarwal

People feel emotions when listening to music. However, emotions are not tangible objects that can be exploited in the music composition process as they are difficult to capture and quantify in algorithms. We present a novel musical…

Artificial Intelligence · Computer Science 2018-10-09 Eunjeong Stella Koh , Shahrokh Yadegari

We present a new system for real-time visualisation of music performance, focused for the moment on a fugue played by a string quartet. The basic principle is to offer a visual guide to better understand music using strategies that should…

Sound · Computer Science 2020-06-19 Olivier Lartillot , Carlos Cancino-Chacón , Charles Brazier

The advent of ML music models such as Google Magenta's MusicVAE now allow us to extract and replicate compositional features from otherwise complex datasets. These models allow computational composers to parameterize abstract variables such…

Multimedia · Computer Science 2021-12-06 Zack Harris , Liam Atticus Clarke , Pietro Gagliano , Dante Camarena , Manal Siddiqui , Pablo S. Castro

Music improvisation is fascinating to study, being essentially a live demonstration of a creative process. In jazz, musicians often improvise across predefined chord progressions (leadsheets). How do we assess the creativity of jazz…

Sound · Computer Science 2025-12-10 Anna Jordanous