English
Related papers

Related papers: Multimodal Dataset Normalization and Perceptual Va…

200 papers

Conditional music generation offers significant advantages in terms of user convenience and control, presenting great potential in AI-generated content research. However, building conditional generative systems for multitrack popular songs…

Sound · Computer Science 2025-10-27 Jing Luo , Xinyu Yang , Dorien Herremans

Prevalent efforts have been put in automatically inferring genres of musical items. Yet, the propose solutions often rely on simplifications and fail to address the diversity and subjectivity of music genres. Accounting for these has,…

Sound · Computer Science 2019-07-30 Elena V. Epure , Anis Khlif , Romain Hennequin

Music generative artificial intelligence (AI) is rapidly expanding music content, necessitating automated song aesthetics evaluation. However, existing studies largely focus on speech, audio or singing quality, leaving song aesthetics…

Sound · Computer Science 2026-01-21 Yishan Lv , Jing Luo , Boyuan Ju , Yang Zhang , Xinda Wu , Bo Yuan , Xinyu Yang

Understanding and modeling consumers' stylistic taste such as "sporty" is crucial for creating designs that truly connect with target audiences. However, capturing taste during the design process remains challenging because taste is…

Human-Computer Interaction · Computer Science 2026-01-27 Matthew K. Hong , Joey Li , Alexandre Filipowicz , Monica Van , Kalani Murakami , Yan-Ying Chen , Shiwali Mohan , Shabnam Hakimi , Matthew Klenk

In our multicultural world, affect-aware AI systems that support humans need the ability to perceive affect across variations in emotion expression patterns across cultures. These systems must perform well in cultural contexts without…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Leena Mathur , Ralph Adolphs , Maja J Matarić

Understanding how auditory stimuli influence emotional and physiological states is fundamental to advancing affective computing and mental health technologies. In this paper, we present a multimodal evaluation of the affective and…

The ability to generate sentiment-controlled feedback in response to multimodal inputs comprising text and images addresses a critical gap in human-computer interaction. This capability allows systems to provide empathetic, accurate, and…

Multimedia · Computer Science 2025-10-07 Puneet Kumar , Sarthak Malik , Balasubramanian Raman , Xiaobai Li

This thesis combines audio-analysis with computer vision to approach Music Information Retrieval (MIR) tasks from a multi-modal perspective. This thesis focuses on the information provided by the visual layer of music videos and how it can…

Multimedia · Computer Science 2020-02-04 Alexander Schindler

This dissertation proposes the study of multimodal learning in the context of musical signals. Throughout, we focus on the interaction between audio signals and text information. Among the many text sources related to music that can be used…

Sound · Computer Science 2021-11-01 Gabriel Meseguer-Brocal

This work combined different audio features to obtain a more robust fingerprint to be used in a music recommendation process. The combination of these methods resulted in a high-dimensional vector. To reduce the number of values, PCA was…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-07 Diego Saldaña Ulloa

Time Scale Modification (TSM) is a well-researched field; however, no effective objective measure of quality exists. This paper details the creation, subjective evaluation, and analysis of a dataset for use in the development of an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-17 Timothy Roberts , Kuldip K. Paliwal

Motifs often recur in musical works in altered forms, preserving aspects of their identity while undergoing local variation. This paper investigates how such motivic transformations occur within their musical context in symbolic music. To…

Sound · Computer Science 2026-03-30 Ron Taieb , Yoel Greenberg , Barak Sober

Obtaining large, human labelled speech datasets to train models for emotion recognition is a notoriously challenging task, hindered by annotation cost and label ambiguity. In this work, we consider the task of learning embeddings for speech…

Computer Vision and Pattern Recognition · Computer Science 2018-08-17 Samuel Albanie , Arsha Nagrani , Andrea Vedaldi , Andrew Zisserman

This paper introduces the MERIT Dataset, a multimodal (text + image + layout) fully labeled dataset within the context of school reports. Comprising over 400 labels and 33k samples, the MERIT Dataset is a valuable resource for training…

Artificial Intelligence · Computer Science 2026-03-04 I. de Rodrigo , A. Sanchez-Cuadrado , J. Boal , A. J. Lopez-Lopez

Pattern discovery algorithms in the music domain aim to find meaningful components in musical compositions. Over the years, although many algorithms have been developed for pattern discovery in music data, it remains a challenging task. To…

Sound · Computer Science 2020-10-26 Iris Ren , Anja Volk , Wouter Swierstra , Remco C. Veltkamp

In this paper we propose a deep learning method for performing attributed-based music-to-image translation. The proposed method is applied for synthesizing visual stories according to the sentiment expressed by songs. The generated images…

Computer Vision and Pattern Recognition · Computer Science 2019-12-13 Nikolaos Passalis , Stavros Doropoulos

We introduce a novel and interpretable path-based music similarity measure. Our similarity measure assumes that items, such as songs and artists, and information about those items are represented in a knowledge graph. We find paths in the…

Information Retrieval · Computer Science 2021-08-05 Giovanni Gabbolini , Derek Bridge

In this work, we tackle a problem of speech emotion classification. One of the issues in the area of affective computation is that the amount of annotated data is very limited. On the other hand, the number of ways that the same emotion can…

Computation and Language · Computer Science 2018-04-02 Egor Lakomkin , Cornelius Weber , Stefan Wermter

As large language models continue to develop, the feasibility and significance of text-based symbolic music tasks have become increasingly prominent. While symbolic music has been widely used in generation tasks, LLM capabilities in…

Sound · Computer Science 2025-09-30 Jiahao Zhao , Yunjia Li , Wei Li , Kazuyoshi Yoshii

Audio-based music structure analysis (MSA) is an essential task in Music Information Retrieval that remains challenging due to the complexity and variability of musical form. Recent advances highlight the potential of fine-tuning…

Sound · Computer Science 2025-07-21 Yixiao Zhang , Haonan Chen , Ju-Chiang Wang , Jitong Chen