English
Related papers

Related papers: Music2P: A Multi-Modal AI-Driven Tool for Simplify…

200 papers

This pictorial aims to critically consider the nature of text-to-audio and text-to-music generative tools in the context of explainable AI. As a group of experimental musicians and researchers, we are enthusiastic about the creative…

Sound · Computer Science 2024-08-15 Jesse Allison , Drew Farrar , Treya Nash , Carlos Román , Morgan Weeks , Fiona Xue Ju

The field of text-to-image (T2I) generation has made significant progress in recent years, largely driven by advancements in diffusion models. Linguistic control enables effective content creation, but struggles with fine-grained control…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Yanan Sun , Yanchen Liu , Yinhao Tang , Wenjie Pei , Kai Chen

Audio is the main form for the visually impaired to obtain information. In reality, all kinds of visual data always exist, but audio data does not exist in many cases. In order to help the visually impaired people to better perceive the…

Sound · Computer Science 2021-03-19 Hailong Ning , Xiangtao Zheng , Yuan Yuan , Xiaoqiang Lu

Generative Adversarial Networks (GANs) were introduced by Goodfellow in 2014, and since then have become popular for constructing generative artificial intelligence models. However, the drawbacks of such networks are numerous, like their…

Computer Vision and Pattern Recognition · Computer Science 2022-12-12 Felipe Perez Stoppa , Ester Vidaña-Vila , Joan Navarro

Recent advances in generative AI music have resulted in new technologies that are being framed as co-creative tools for musicians with early work demonstrating their potential to add to music practice. While the field has seen many valuable…

Human-Computer Interaction · Computer Science 2025-02-14 Stephen James Krol , Maria Teresa Llano Rodriguez , Miguel Loor Paredes

Automatic music captioning, which generates natural language descriptions for given music tracks, holds significant potential for enhancing the understanding and organization of large volumes of musical data. Despite its importance,…

Sound · Computer Science 2023-08-01 SeungHeon Doh , Keunwoo Choi , Jongpil Lee , Juhan Nam

The integration of artificial intelligence (AI) technology in the music industry is driving a significant change in the way music is being composed, produced and mixed. This study investigates the current state of AI in the mixing workflows…

Human-Computer Interaction · Computer Science 2023-09-11 Soumya Sai Vanka , Maryam Safi , Jean-Baptiste Rolland , George Fazekas

Recent work in the field of symbolic music generation has shown value in using a tokenization based on the GuitarPro format, a symbolic representation supporting guitar expressive attributes, as an input and output representation. We extend…

Sound · Computer Science 2023-07-12 Jackson Loth , Pedro Sarmento , CJ Carr , Zack Zukowski , Mathieu Barthet

Music is a powerful medium for altering the emotional state of the listener. In recent years, with significant advancement in computing capabilities, artificial intelligence-based (AI-based) approaches have become popular for creating…

Human-Computer Interaction · Computer Science 2023-01-18 Adyasha Dash , Kat R. Agres

Owing to the power of vision-language foundation models, e.g., CLIP, the area of image synthesis has seen recent important advances. Particularly, for style transfer, CLIP enables transferring more general and abstract styles without…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Zipeng Xu , Songlong Xing , Enver Sangineto , Nicu Sebe

Embodied agents have achieved prominent performance in following human instructions to complete tasks. However, the potential of providing instructions informed by texts and images to assist humans in completing tasks remains underexplored.…

Computation and Language · Computer Science 2023-05-04 Yujie Lu , Pan Lu , Zhiyu Chen , Wanrong Zhu , Xin Eric Wang , William Yang Wang

This paper introduces text2midi, an end-to-end model to generate MIDI files from textual descriptions. Leveraging the growing popularity of multimodal generative approaches, text2midi capitalizes on the extensive availability of textual…

Sound · Computer Science 2025-06-18 Keshav Bhandari , Abhinaba Roy , Kyra Wang , Geeta Puri , Simon Colton , Dorien Herremans

Fashion is a fast-changing industry where designs are refreshed at large scale every season. Moreover, it faces huge challenge of unsold inventory as not all designs appeal to customers. This puts designers under significant pressure.…

Computer Vision and Pattern Recognition · Computer Science 2020-07-13 Alpana Dubey , Nitish Bhardwaj , Kumar Abhinav , Suma Mani Kuriakose , Sakshi Jain , Veenu Arora

Designed light is an established modality for live performance and music playback. Despite the growing availability of consumer smart lighting, the creation of designed light for music visualization remains limited to professional contexts…

Human-Computer Interaction · Computer Science 2026-02-10 Frederic Anthony Robinson , Vishnu Raj , David Cooper , Fan Du , David Gunawan

The unabated mystique of large-scale neural networks, such as the CLIP dual image-and-text encoder, popularized automatically generated art. Increasingly more sophisticated generators enhanced the artworks' realism and visual appearance,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Piotr Mirowski , Dylan Banarse , Mateusz Malinowski , Simon Osindero , Chrisantha Fernando

While many tools are available for designing AI, non-experts still face challenges in clearly expressing their intent and managing system complexity. We introduce AIAP, a no-code platform that integrates natural language input with visual…

Human-Computer Interaction · Computer Science 2025-08-05 Hyunjn An , Yongwon Kim , Wonduk Seo , Joonil Park , Daye Kang , Changhoon Oh , Dokyun Kim , Seunghyun Lee

AI illustrator aims to automatically design visually appealing images for books to provoke rich thoughts and emotions. To achieve this goal, we propose a framework for translating raw descriptions with complex semantics into semantically…

Computer Vision and Pattern Recognition · Computer Science 2022-09-09 Yiyang Ma , Huan Yang , Bei Liu , Jianlong Fu , Jiaying Liu

Content-based music information retrieval has seen rapid progress with the adoption of deep learning. Current approaches to high-level music description typically make use of classification models, such as in auto-tagging or genre and mood…

Sound · Computer Science 2021-12-09 Ilaria Manco , Emmanouil Benetos , Elio Quinton , Gyorgy Fazekas

The burgeoning growth of video-to-music generation can be attributed to the ascendancy of multimodal generative models. However, there is a lack of literature that comprehensively combs through the work in this field. To fill this gap, this…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-23 Shulei Ji , Songruoyao Wu , Zihao Wang , Shuyu Li , Kejun Zhang

Generating images that fit a given text description using machine learning has improved greatly with the release of technologies such as the CLIP image-text encoder model; however, current methods lack artistic control of the style of image…

Computer Vision and Pattern Recognition · Computer Science 2022-03-02 Peter Schaldenbrand , Zhixuan Liu , Jean Oh