中文
相关论文

相关论文: Music2P: A Multi-Modal AI-Driven Tool for Simplify…

200 篇论文

We propose a system that learns from artistic pairings of music and corresponding album cover art. The goal is to 'translate' paintings into music and, in further stages of development, the converse. We aim to deploy this system as an…

声音 · 计算机科学 2020-08-25 Prateek Verma , Constantin Basica , Pamela Davis Kivelson

Existing work in automatic music generation has mostly focused on end-to-end systems that generate either entire compositions or continuations of pieces, which are difficult for composers to iterate on. The area of computer-assisted…

声音 · 计算机科学 2026-01-27 Christian Zhou-Zheng , Philippe Pasquier

Image-to-point-cloud (I2P) registration aims to align 2D images with 3D point clouds by establishing reliable 2D-3D correspondences. The drastic modality gap between images and point clouds makes it challenging to learn features that are…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Pei An , Junfeng Ding , Jiaqi Yang , Yulong Wang , Jie Ma , Liangliang Nan

Cover songs constitute a vital aspect of musical culture, preserving the core melody of an original composition while reinterpreting it to infuse novel emotional depth and thematic emphasis. Although prior research has explored the…

声音 · 计算机科学 2026-02-24 Sifei Li , Yang Li , Zizhou Wang , Yuxin Zhang , Fuzhang Wu , Oliver Deussen , Tong-Yee Lee , Weiming Dong

AI tools increasingly shape how we discover, make and experience music. While these tools can have the potential to empower creativity, they may fundamentally redefine relationships between stakeholders, to the benefit of some and the…

Creating MIDI music can be a practical challenge. In the past, working with it was difficult and frustrating to all but the most accomplished and determined. Now, however, we are offering a powerful Visual Basic program called MIDI-LAB,…

软件工程 · 计算机科学 2015-03-13 Kai Yang , Xi Zhou

AI systems for high quality music generation typically rely on extremely large musical datasets to train the AI models. This creates barriers to generating music beyond the genres represented in dominant datasets such as Western Classical…

声音 · 计算机科学 2024-07-19 Nick Bryan-Kinns , Zijin Li

Content creators often draw inspiration from multiple visual sources, combining distinct elements to craft new compositions. Modern computational approaches now aim to emulate this fundamental creative process. Although recent diffusion…

计算机视觉与模式识别 · 计算机科学 2025-02-20 Sara Dorfman , Dana Cohen-Bar , Rinon Gal , Daniel Cohen-Or

Computer-generated visualisations can accompany recorded or live music to create novel audiovisual experiences for audiences. We present a system to streamline the creation of audio-driven visualisations based on audio feature extraction…

多媒体 · 计算机科学 2021-06-21 Max Graf , Harold Chijioke Opara , Mathieu Barthet

Multi-modal music generation, using multiple modalities like text, images, and video alongside musical scores and audio as guidance, is an emerging research area with broad applications. This paper reviews this field, categorizing music…

声音 · 计算机科学 2026-03-09 Shuyu Li , Shulei Ji , Zihao Wang , Songruoyao Wu , Jiaxing Yu , Kejun Zhang

Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist due to the lack of a captioned MIDI dataset. This work aims to enable research that combines LLMs with symbolic…

音频与语音处理 · 电气工程与系统科学 2025-08-08 Jan Melechovsky , Abhinaba Roy , Dorien Herremans

Computational aesthetic evaluation has made remarkable contribution to visual art works, but its application to music is still rare. Currently, subjective evaluation is still the most effective form of evaluating artistic works. However,…

计算机视觉与模式识别 · 计算机科学 2024-02-14 Xin Jin , Wu Zhou , Jingyu Wang , Duo Xu , Yongsen Zheng

In the task of generating music, the art factor plays a big role and is a great challenge for AI. Previous work involving adversarial training to produce new music pieces and modeling the compatibility of variety in music (beats, tempo,…

声音 · 计算机科学 2023-01-09 Abhinav Kaushal Keshari

The burgeoning volume of digital content across diverse modalities necessitates efficient storage and retrieval methods. Conventional approaches struggle to cope with the escalating complexity and scale of multimedia data. In this paper, we…

人工智能 · 计算机科学 2024-04-17 Jixiang Luo

Songwriting has long served as a powerful medium for expressing unconscious emotions and fostering self-awareness in psychotherapy. Due to the auditory-centric nature of traditional approaches, Deaf and Hard-of-Hearing (DHH) individuals…

人机交互 · 计算机科学 2026-04-16 Youjin Choi , Jaeyoung Moon , Jinyoung Yoo , Jennifer G. Kim , Jin-Hyuk Hong

Recent research in the field of computer vision strongly focuses on deep learning architectures to tackle image processing problems. Deep neural networks are often considered in complex image processing scenarios since traditional computer…

计算机视觉与模式识别 · 计算机科学 2021-11-30 Marcel P. Schilling , Luca Rettenberger , Friedrich Münke , Haijun Cui , Anna A. Popova , Pavel A. Levkin , Ralf Mikut , Markus Reischl

Generating a complex work of art such as a musical composition requires exhibiting true creativity that depends on a variety of factors that are related to the hierarchy of musical language. Music generation have been faced with Algorithmic…

声音 · 计算机科学 2021-09-08 Carlos Hernandez-Olivan , Jose R. Beltran

With the rise of artificial intelligence in recent years, there has been a rapid increase in its application towards creative domains, including music. There exist many systems built that apply machine learning approaches to the problem of…

人机交互 · 计算机科学 2025-04-22 Renaud Bougueng Tchemeube , Jeff Ens , Philippe Pasquier

Learning cross-modal correspondences is essential for image-to-point cloud (I2P) registration. Existing methods achieve this mostly by utilizing metric learning to enforce feature alignment across modalities, disregarding the inherent…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Juncheng Mu , Chengwei Ren , Weixiang Zhang , Liang Pan , Xiao-Ping Zhang , Yue Gao

This paper introduces the WordArt Designer API, a novel framework for user-driven artistic typography synthesis utilizing Large Language Models (LLMs) on ModelScope. We address the challenge of simplifying artistic typography for…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Jun-Yan He , Zhi-Qi Cheng , Chenyang Li , Jingdong Sun , Wangmeng Xiang , Yusen Hu , Xianhui Lin , Xiaoyang Kang , Zengke Jin , Bin Luo , Yifeng Geng , Xuansong Xie , Jingren Zhou