中文
相关论文

相关论文: Cross-Modal Learning for Music-to-Music-Video Desc…

200 篇论文

The burgeoning growth of video-to-music generation can be attributed to the ascendancy of multimodal generative models. However, there is a lack of literature that comprehensively combs through the work in this field. To fill this gap, this…

音频与语音处理 · 电气工程与系统科学 2025-12-23 Shulei Ji , Songruoyao Wu , Zihao Wang , Shuyu Li , Kejun Zhang

Multi-modal music generation, using multiple modalities like text, images, and video alongside musical scores and audio as guidance, is an emerging research area with broad applications. This paper reviews this field, categorizing music…

声音 · 计算机科学 2026-03-09 Shuyu Li , Shulei Ji , Zihao Wang , Songruoyao Wu , Jiaxing Yu , Kejun Zhang

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Zhaokai Wang , Chenxi Bao , Le Zhuo , Jingrui Han , Yang Yue , Yihong Tang , Victor Shea-Jay Huang , Yue Liao

Multimodal music generation aims to produce music from diverse input modalities, including text, videos, and images. Existing methods use a common embedding space for multimodal fusion. Despite their effectiveness in other modalities, their…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Baisen Wang , Le Zhuo , Zhaokai Wang , Chenxi Bao , Wu Chengjing , Xuecheng Nie , Jiao Dai , Jizhong Han , Yue Liao , Si Liu

Video-to-music generation presents significant potential in video production, requiring the generated music to be both semantically and rhythmically aligned with the video. Achieving this alignment demands advanced music generation…

声音 · 计算机科学 2024-12-10 Sifei Li , Binxin Yang , Chunji Yin , Chong Sun , Yuxin Zhang , Weiming Dong , Chen Li

Creation of images using generative adversarial networks has been widely adapted into multi-modal regime with the advent of multi-modal representation models pre-trained on large corpus. Various modalities sharing a common representation…

声音 · 计算机科学 2022-06-10 Yoonjeon Kim , Joel Jang , Sumin Shin

In this study, we explore the representation mapping from the domain of visual arts to the domain of music, with which we can use visual arts as an effective handle to control music generation. Unlike most studies in multimodal…

声音 · 计算机科学 2022-11-11 Runbang Zhang , Yixiao Zhang , Kai Shao , Ying Shan , Gus Xia

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced audio codecs, the…

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box…

多媒体 · 计算机科学 2025-07-29 Junxian Wu , Weitao You , Heda Zuo , Dengming Zhang , Pei Chen , Lingyun Sun

We present the Melody-Guided Music Generation (MG2) model, a novel approach using melody to guide the text-to-music generation that, despite a simple method and limited resources, achieves excellent performance. Specifically, we first align…

声音 · 计算机科学 2024-12-31 Shaopeng Wei , Manzhen Wei , Haoyu Wang , Yu Zhao , Gang Kou

Cross-modal audio-visual perception has been a long-lasting topic in psychology and neurology, and various studies have discovered strong correlations in human perception of auditory and visual stimuli. Despite works in computational…

计算机视觉与模式识别 · 计算机科学 2017-04-28 Lele Chen , Sudhanshu Srivastava , Zhiyao Duan , Chenliang Xu

Numerous studies in the field of music generation have demonstrated impressive performance, yet virtually no models are able to directly generate music to match accompanying videos. In this work, we develop a generative music AI framework,…

声音 · 计算机科学 2024-06-03 Jaeyong Kang , Soujanya Poria , Dorien Herremans

Text-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this process demands extensive, high-quality data, which are often…

声音 · 计算机科学 2025-06-18 Chang Li , Ruoyu Wang , Lijuan Liu , Jun Du , Yixuan Sun , Zilu Guo , Zhenrong Zhang , Yuan Jiang , Jianqing Gao , Feng Ma

A major challenge in text-video and text-audio retrieval is the lack of large-scale training data. This is unlike image-captioning, where datasets are in the order of millions of samples. To close this gap we propose a new video mining…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Arsha Nagrani , Paul Hongsuck Seo , Bryan Seybold , Anja Hauth , Santiago Manen , Chen Sun , Cordelia Schmid

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…

In this work, we study music/video cross-modal recommendation, i.e. recommending a music track for a video or vice versa. We rely on a self-supervised learning paradigm to learn from a large amount of unlabelled data. We rely on a…

多媒体 · 计算机科学 2021-05-03 Laure Pretet , Gael Richard , Geoffroy Peeters

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

Music generation aims to create music segments that align with human aesthetics based on diverse conditional information. Despite advancements in generating music from specific textual descriptions (e.g., style, genre, instruments), the…

声音 · 计算机科学 2025-04-21 Jiahao Song , Yuzhao Wang

Having access to multi-modal cues (e.g. vision and audio) empowers some cognitive tasks to be done faster compared to learning from a single modality. In this work, we propose to transfer knowledge across heterogeneous modalities, even…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Yanbei Chen , Yongqin Xian , A. Sophia Koepke , Ying Shan , Zeynep Akata
‹ 上一页 1 2 3 10 下一页 ›