中文
相关论文

相关论文: From Sound to Sight: Towards AI-authored Music Vid…

200 篇论文

The burgeoning growth of video-to-music generation can be attributed to the ascendancy of multimodal generative models. However, there is a lack of literature that comprehensively combs through the work in this field. To fill this gap, this…

音频与语音处理 · 电气工程与系统科学 2025-12-23 Shulei Ji , Songruoyao Wu , Zihao Wang , Shuyu Li , Kejun Zhang

Visuals can enhance our experience of music, owing to the way they can amplify the emotions and messages conveyed within it. However, creating music visualization is a complex, time-consuming, and resource-intensive process. We introduce…

人机交互 · 计算机科学 2023-09-29 Vivian Liu , Tao Long , Nathan Raw , Lydia Chilton

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Zhaokai Wang , Chenxi Bao , Le Zhuo , Jingrui Han , Yang Yue , Yihong Tang , Victor Shea-Jay Huang , Yue Liao

Video-to-music generation presents significant potential in video production, requiring the generated music to be both semantically and rhythmically aligned with the video. Achieving this alignment demands advanced music generation…

声音 · 计算机科学 2024-12-10 Sifei Li , Binxin Yang , Chunji Yin , Chong Sun , Yuxin Zhang , Weiming Dong , Chen Li

In this paper we propose a deep learning method for performing attributed-based music-to-image translation. The proposed method is applied for synthesizing visual stories according to the sentiment expressed by songs. The generated images…

计算机视觉与模式识别 · 计算机科学 2019-12-13 Nikolaos Passalis , Stavros Doropoulos

Creation of images using generative adversarial networks has been widely adapted into multi-modal regime with the advent of multi-modal representation models pre-trained on large corpus. Various modalities sharing a common representation…

声音 · 计算机科学 2022-06-10 Yoonjeon Kim , Joel Jang , Sumin Shin

This paper is a survey and an analysis of different ways of using deep learning (deep artificial neural networks) to generate musical content. We propose a methodology based on five dimensions for our analysis: Objective - What musical…

声音 · 计算机科学 2019-08-09 Jean-Pierre Briot , Gaëtan Hadjeres , François-David Pachet

In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments. We first identify two key intermediate representations for a successful video to music…

计算机视觉与模式识别 · 计算机科学 2020-07-22 Chuang Gan , Deng Huang , Peihao Chen , Joshua B. Tenenbaum , Antonio Torralba

Generating a complex work of art such as a musical composition requires exhibiting true creativity that depends on a variety of factors that are related to the hierarchy of musical language. Music generation have been faced with Algorithmic…

声音 · 计算机科学 2021-09-08 Carlos Hernandez-Olivan , Jose R. Beltran

Deep Learning models have shown very promising results in automatically composing polyphonic music pieces. However, it is very hard to control such models in order to guide the compositions towards a desired goal. We are interested in…

机器学习 · 计算机科学 2021-03-11 Lucas N. Ferreira , Jim Whitehead

As two of the five traditional human senses (sight, hearing, taste, smell, and touch), vision and sound are basic sources through which humans understand the world. Often correlated during natural events, these two modalities combine to…

计算机视觉与模式识别 · 计算机科学 2018-06-04 Yipin Zhou , Zhaowen Wang , Chen Fang , Trung Bui , Tamara L. Berg

In this work, we provide a comprehensive survey of AI music generation tools, including both research projects and commercialized applications. To conduct our analysis, we classified music generation approaches into three categories:…

声音 · 计算机科学 2023-08-28 Yueyue Zhu , Jared Baca , Banafsheh Rekabdar , Reza Rawassizadeh

Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generating music that…

声音 · 计算机科学 2024-09-13 Tanisha Hisariya , Huan Zhang , Jinhua Liang

How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with the large gaps that often exist between sight and sound. We…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Andrew Owens , Tae-Hyun Oh

The utilization of deep learning techniques in generating various contents (such as image, text, etc.) has become a trend. Especially music, the topic of this paper, has attracted widespread attention of countless researchers.The whole…

声音 · 计算机科学 2020-11-16 Shulei Ji , Jing Luo , Xinyu Yang

This study presents a novel method for generating music visualisers using diffusion models, combining audio input with user-selected artwork. The process involves two main stages: image generation and video creation. First, music captioning…

多媒体 · 计算机科学 2024-12-10 Leonardo Pina , Yongmin Li

Artificial Intelligence and generative models have revolutionized music creation, with many models leveraging textual or visual prompts for guidance. However, existing image-to-music models are limited to simple images, lacking the…

多媒体 · 计算机科学 2025-07-31 Ivan Rinaldi , Nicola Fanelli , Giovanna Castellano , Gennaro Vessio

In addition to traditional tasks such as prediction, classification and translation, deep learning is receiving growing attention as an approach for music generation, as witnessed by recent research groups such as Magenta at Google and CTRL…

声音 · 计算机科学 2018-11-13 Jean-Pierre Briot , François Pachet

Generative AI has been transforming the way we interact with technology and consume content. In the next decade, AI technology will reshape how we create audio content in various media, including music, theater, films, games, podcasts, and…

声音 · 计算机科学 2024-11-25 Hao-Wen Dong

As Artificial Intelligence (AI) technologies continue to evolve, their use in generating realistic, contextually appropriate content has expanded into various domains. Music, an art form and medium for entertainment, deeply rooted into…

声音 · 计算机科学 2024-12-11 Yupei Li , Manuel Milling , Lucia Specia , Björn W. Schuller
‹ 上一页 1 2 3 10 下一页 ›