English
Related papers

Related papers: Emotion-Guided Image to Music Generation

200 papers

Deep generative models such as Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) have recently been applied to style and domain transfer for images, and in the case of VAEs, music. GAN-based models employing several…

Sound · Computer Science 2018-09-21 Gino Brunner , Yuyi Wang , Roger Wattenhofer , Sumu Zhao

Introduction: Music provides an incredible avenue for individuals to express their thoughts and emotions, while also serving as a delightful mode of entertainment for enthusiasts and music lovers. Objectives: This paper presents a…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Rajesh B , Keerthana V , Narayana Darapaneni , Anwesh Reddy P

Convolutional Neural Networks are particularly suited for image analysis tasks, such as Image Classification, Object Recognition or Image Segmentation. Like all Artificial Neural Networks, however, they are "black box" models, and suffer…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Youssef Doulfoukar , Laurent Mertens , Joost Vennekens

Multi-modal music generation, using multiple modalities like text, images, and video alongside musical scores and audio as guidance, is an emerging research area with broad applications. This paper reviews this field, categorizing music…

Sound · Computer Science 2026-03-09 Shuyu Li , Shulei Ji , Zihao Wang , Songruoyao Wu , Jiaxing Yu , Kejun Zhang

An image conveys meaning through both its visual content and emotional tone, jointly shaping human perception. We introduce Controllable Emotional Image Content Generation (C-EICG), which aims to generate images that remain faithful to a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Jingyuan Yang , Weibin Luo , Hui Huang

This paper addresses the problem of sheet-image-based on-line audio-to-score alignment also known as score following. Drawing inspiration from object detection, a conditional neural network architecture is proposed that directly predicts…

Sound · Computer Science 2021-05-11 Florian Henkel , Gerhard Widmer

In this paper, we propose a cross-modal variational auto-encoder (CMVAE) for content-based micro-video background music recommendation. CMVAE is a hierarchical Bayesian generative model that matches relevant background music to a…

Multimedia · Computer Science 2022-12-13 Jing Yi , Yaochen Zhu , Jiayi Xie , Zhenzhong Chen

We present in this paper PerformacnceNet, a neural network model we proposed recently to achieve score-to-audio music generation. The model learns to convert a music piece from the symbolic domain to the audio domain, assigning…

Sound · Computer Science 2019-05-29 Yu-Hua Chen , Bryan Wang , Yi-Hsuan Yang

Transformers and variational autoencoders (VAE) have been extensively employed for symbolic (e.g., MIDI) domain music generation. While the former boast an impressive capability in modeling long sequences, the latter allow users to…

Sound · Computer Science 2022-12-21 Shih-Lun Wu , Yi-Hsuan Yang

In this study, we aim to determine if generalized sounds and music can share a common emotional space, improving predictions of emotion in terms of arousal and valence. We propose the use of multiple datasets as a multi-domain learning…

Sound · Computer Science 2024-08-15 Federico Simonetta , Francesca Certo , Stavros Ntalampiras

The field of automatic music composition has seen great progress in the last few years, much of which can be attributed to advances in deep neural networks. There are numerous studies that present different strategies for generating sheet…

Sound · Computer Science 2021-04-28 Dimos Makris , Kat R. Agres , Dorien Herremans

Large text-guided diffusion models, such as DALLE-2, are able to generate stunning photorealistic images given natural language descriptions. While such models are highly flexible, they struggle to understand the composition of certain…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Nan Liu , Shuang Li , Yilun Du , Antonio Torralba , Joshua B. Tenenbaum

Synthesizing realistic data samples is of great value for both academic and industrial communities. Deep generative models have become an emerging topic in various research areas like computer vision and signal processing. Affective…

Computer Vision and Pattern Recognition · Computer Science 2020-11-10 Noushin Hajarolasvadi , Miguel Arjona Ramírez , Hasan Demirel

The Synesthetic Variational Autoencoder (SynVAE) introduced in this research is able to learn a consistent mapping between visual and auditive sensory modalities in the absence of paired datasets. A quantitative evaluation on MNIST as well…

Computer Vision and Pattern Recognition · Computer Science 2019-09-15 Maximilian Müller-Eberstein , Nanne van Noord

In the domain of human-computer interaction, accurately recognizing and interpreting human emotions is crucial yet challenging due to the complexity and subtlety of emotional expressions. This study explores the potential for detecting a…

Multimedia · Computer Science 2025-05-13 Jiehui Jia , Huan Zhang , Jinhua Liang

Recent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow like motion feature representations, which exhibit limited…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Chuang Gan , Deng Huang , Hang Zhao , Joshua B. Tenenbaum , Antonio Torralba

The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Seung Hyun Lee , Gyeongrok Oh , Wonmin Byeon , Chanyoung Kim , Won Jeong Ryoo , Sang Ho Yoon , Hyunjun Cho , Jihyun Bae , Jinkyu Kim , Sangpil Kim

Multimodal emotion recognition has recently gained much attention since it can leverage diverse and complementary relationships over multiple modalities (e.g., audio, visual, biosignals, etc.), and can provide some robustness to noisy…

This paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Zehuan Huang , Yuan-Chen Guo , Xingqiao An , Yunhan Yang , Yangguang Li , Zi-Xin Zou , Ding Liang , Xihui Liu , Yan-Pei Cao , Lu Sheng

We present a deep metric variational autoencoder for multi-modal data generation. The variational autoencoder employs triplet loss in the latent space, which allows for conditional data generation by sampling in the latent space within each…