English
Related papers

Related papers: Emotion-Conditioned Melody Harmonization with Hier…

200 papers

A video prediction model that generalizes to diverse scenes would enable intelligent agents such as robots to perform a variety of tasks via planning with the model. However, while existing video prediction models have produced promising…

Computer Vision and Pattern Recognition · Computer Science 2021-06-22 Bohan Wu , Suraj Nair , Roberto Martin-Martin , Li Fei-Fei , Chelsea Finn

Re-orchestration is the process of adapting a music piece for a different set of instruments. By altering the original instrumentation, the orchestrator often modifies the musical texture while preserving a recognizable melodic line and…

Sound · Computer Science 2025-07-01 Dinh-Viet-Toan Le , Yi-Hsuan Yang

In human interactions, emotion recognition is crucial. For this reason, the topic of computer-vision approaches for automatic emotion recognition is currently being extensively researched. Processing multi-channel electroencephalogram (EEG)…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Joshua Bègue , Mohamed Aymen Labiod , Abdelhamid Melloulk

Vision-language models (VLMs) have shown impressive abilities across a range of multi-modal tasks. However, existing metrics for evaluating the quality of text generated by VLMs typically focus on an overall evaluation for a specific task,…

Computation and Language · Computer Science 2026-03-10 Masanari Ohi , Masahiro Kaneko , Naoaki Okazaki , Nakamasa Inoue

Variational Autoencoders (VAEs) have proven to be effective models for producing latent representations of cognitive and semantic value. We assess the degree to which VAEs trained on a prototypical tonal music corpus of 371 Bach's chorales…

Sound · Computer Science 2023-11-08 Nádia Carvalho , Gilberto Bernardes

Tracking an interpretable emotional arc of a conversation via the sentiment of individual utterances processed as a whole is central to both understanding and guiding communication in applied, especially clinical, conversational contexts.…

Artificial Intelligence · Computer Science 2026-05-14 Anamika Ragu , Aneesh Jonelagadda

Learning latent representations that are simultaneously expressive, geometrically well-structured, and reliably calibrated remains a central challenge for Variational Autoencoders (VAEs). Standard VAEs typically assume a diagonal Gaussian…

Machine Learning · Computer Science 2025-12-02 Mehmet Can Yavuz

The rise of deep learning technologies has quickly advanced many fields, including that of generative music systems. There exist a number of systems that allow for the generation of good sounding short snippets, yet, these generated…

Sound · Computer Science 2021-04-27 Zixun Guo , Makris Dimos , Herremans Dorien

While generative models have shown great success in generating high-dimensional samples conditional on low-dimensional descriptors (learning e.g. stroke thickness in MNIST, hair color in CelebA, or speaker identity in Wavenet), their…

Machine Learning · Computer Science 2019-10-31 Mohammad Lotfollahi , Mohsen Naghipourfar , Fabian J. Theis , F. Alexander Wolf

Emotion alignment between music and palettes is crucial for effective multimedia content, yet misalignment creates confusion that weakens the intended message. However, existing methods often generate only a single dominant color, missing…

Multimedia · Computer Science 2025-09-18 Jiayun Hu , Yueyi He , Tianyi Liang , Changbo Wang , Chenhui Li

We address the challenging open problem of learning an effective latent space for symbolic music data in generative music modeling. We focus on leveraging adversarial regularization as a flexible and natural mean to imbue variational…

Sound · Computer Science 2020-02-21 Andrea Valenti , Antonio Carta , Davide Bacciu

Predicting customers' long-term revenue from sparse and irregular transaction data is central to marketing resource allocation in non-contractual settings, yet existing approaches face a trade-off. Traditional probabilistic customer base…

Machine Learning · Statistics 2026-04-27 Jeffrey Näf , Riana Valera Mbelson , Markus Meierer

Video-based Affective Computing (VAC), vital for emotion analysis and human-computer interaction, suffers from model instability and representational degradation due to complex emotional dynamics. Since the meaning of different emotional…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Feng-Qi Cui , Jinyang Huang , Sirui Zhao , Xinyu Li , Xin Yan , Ziyu Jia , Xiaokang Zhou

Discovering and exploring the underlying structure of multi-instrumental music using learning-based approaches remains an open problem. We extend the recent MusicVAE model to represent multitrack polyphonic measures as vectors in a latent…

Machine Learning · Statistics 2018-06-04 Ian Simon , Adam Roberts , Colin Raffel , Jesse Engel , Curtis Hawthorne , Douglas Eck

This paper proposes a new model for music prediction based on Variational Autoencoders (VAEs). In this work, VAEs are used in a novel way in order to address two different problems: music representation into the latent space, and using this…

Sound · Computer Science 2019-06-25 Daniel Rivero , Enrique Fernandez-Blanco , Alejandro Pazos

Variational Autoencoders(VAEs) have already achieved great results on image generation and recently made promising progress on music generation. However, the generation process is still quite difficult to control in the sense that the…

Sound · Computer Science 2019-04-19 Ruihan Yang , Tianyao Chen , Yiyi Zhang , Gus Xia

The field of automatic music composition has seen great progress in the last few years, much of which can be attributed to advances in deep neural networks. There are numerous studies that present different strategies for generating sheet…

Sound · Computer Science 2021-04-28 Dimos Makris , Kat R. Agres , Dorien Herremans

The task of predicting stochastic behaviors of road agents in diverse environments is a challenging problem for autonomous driving. To best understand scene contexts and produce diverse possible future states of the road agents adaptively…

Machine Learning · Computer Science 2022-01-25 Geunseob Oh , Huei Peng

In this work we study Variational Autoencoders (VAEs) from the perspective of harmonic analysis. By viewing a VAE's latent space as a Gaussian Space, a variety of measure space, we derive a series of results that show that the encoder…

Machine Learning · Statistics 2022-04-26 Alexander Camuto , Matthew Willetts

Multimodal music emotion analysis leverages both audio and MIDI modalities to enhance performance. While mainstream approaches focus on complex feature extraction networks, we propose that shortening the length of audio sequence features to…

Sound · Computer Science 2025-09-24 Dinghao Zou , Yicheng Gong , Xiaokang Li , Xin Cao , Sunbowen Lee