English

What Do I Hear? Generating Sounds for Visuals with ChatGPT

Sound 2023-11-10 v1 Computer Vision and Pattern Recognition Multimedia Audio and Speech Processing

Abstract

This short paper introduces a workflow for generating realistic soundscapes for visual media. In contrast to prior work, which primarily focus on matching sounds for on-screen visuals, our approach extends to suggesting sounds that may not be immediately visible but are essential to crafting a convincing and immersive auditory environment. Our key insight is leveraging the reasoning capabilities of language models, such as ChatGPT. In this paper, we describe our workflow, which includes creating a scene context, brainstorming sounds, and generating the sounds.

Keywords

Cite

@article{arxiv.2311.05609,
  title  = {What Do I Hear? Generating Sounds for Visuals with ChatGPT},
  author = {David Chuan-En Lin and Nikolas Martelaro},
  journal= {arXiv preprint arXiv:2311.05609},
  year   = {2023}
}

Comments

Demo: http://soundify.cc

R2 v1 2026-06-28T13:16:38.881Z