English
Related papers

Related papers: What Do I Hear? Generating Sounds for Visuals with…

200 papers

Gesture synthesis has gained significant attention as a critical research field, aiming to produce contextually appropriate and natural gestures corresponding to speech or textual input. Although deep learning-based approaches have achieved…

Computation and Language · Computer Science 2024-05-29 Nan Gao , Zeyu Zhao , Zhi Zeng , Shuwu Zhang , Dongdong Weng , Yihua Bao

This work proposes a novel method to generate realistic talking head videos using audio and visual streams. We animate a source image by transferring head motion from a driving video using a dense motion field generated using learnable…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Madhav Agarwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

In this paper, we present MovieFactory, a powerful framework to generate cinematic-picture (3072$\times$1280), film-style (multi-scene), and multi-modality (sounding) movies on the demand of natural languages. As the first fully automated…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Junchen Zhu , Huan Yang , Huiguo He , Wenjing Wang , Zixi Tuo , Wen-Huang Cheng , Lianli Gao , Jingkuan Song , Jianlong Fu

The content of visual and audio scenes is multi-faceted such that a video can be paired with various audio and vice-versa. Thereby, in video-to-audio generation task, it is imperative to introduce steering approaches for controlling the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Xiulong Liu , Kun Su , Eli Shlizerman

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic videos are disregarded.…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Jiawei Liu , Weining Wang , Sihan Chen , Xinxin Zhu , Jing Liu

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the recent success, current LLMs are not capable of processing…

Launchpad is a musical instrument that allows users to create and perform music by pressing illuminated buttons. To assist and inspire the design of the Launchpad light effect, and provide a more accessible approach for beginners to create…

Sound · Computer Science 2025-10-09 Siting Xu , Yolo Yunlong Tang , Feng Zheng

Language models pretrained on text-only corpora often struggle with tasks that require auditory commonsense knowledge. Previous work addresses this problem by augmenting the language model to retrieve knowledge from external audio…

Computation and Language · Computer Science 2025-06-10 Suho Yoo , Hyunjong Ok , Jaeho Lee

ChatGPT is a natural language processing tool that can engage in human-like conversations and generate coherent and contextually relevant responses to various prompts. ChatGPT is capable of understanding natural text that is input by a user…

Human-Computer Interaction · Computer Science 2024-09-02 Sara Amani , Lance White , Trini Balart , Laksha Arora , Kristi J. Shryock , Kelly Brumbelow , Karan L. Watson

Synthesizing realistic images from text descriptions on a dataset like Microsoft Common Objects in Context (MS COCO), where each image can contain several objects, is a challenging task. Prior work has used text captions to generate images.…

Computer Vision and Pattern Recognition · Computer Science 2018-02-23 Shikhar Sharma , Dendi Suhubdy , Vincent Michalski , Samira Ebrahimi Kahou , Yoshua Bengio

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-shot settings. To tackle the challenge of limited labeled…

Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene Synthesis challenge held as part of the Detection and…

We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try…

Graphics · Computer Science 2024-09-23 Matthew Caren , Kartik Chandra , Joshua B. Tenenbaum , Jonathan Ragan-Kelley , Karima Ma

In this contribution, we will discuss a prototype that allows a group of users to design sound collaboratively in real time using a multi-touch tabletop. We make use of a machine learning method to generate a mapping from perceptual audio…

Multimedia · Computer Science 2014-06-24 Niklas Klügel , Timo Becker , Georg Groh

Objects produce different sounds when hit, and humans can intuitively infer how an object might sound based on its appearance and material properties. Inspired by this intuition, we propose Visual Acoustic Fields, a framework that bridges…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Yuelei Li , Hyunjin Kim , Fangneng Zhan , Ri-Zhao Qiu , Mazeyu Ji , Xiaojun Shan , Xueyan Zou , Paul Liang , Hanspeter Pfister , Xiaolong Wang

The importance of recommender systems on the web has grown, especially in the movie industry, with a vast selection of options to watch. To assist users in traversing available items and finding relevant results, recommender systems analyze…

Information Retrieval · Computer Science 2025-07-30 Ali Fallahi , Azam Bastanfard , Amineh Amini , Hadi Saboohi

With the introduction of ChatGPT, the public's perception of AI-generated content (AIGC) has begun to reshape. Artificial intelligence has significantly reduced the barrier to entry for non-professionals in creative endeavors, enhancing the…

Sound · Computer Science 2023-11-21 Lei Wang , Ziyi Zhao , Hanwei Liu , Junwei Pang , Yi Qin , Qidi Wu

Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have…

Sound · Computer Science 2024-10-17 Mithun Manivannan , Vignesh Nethrapalli , Mark Cartwright

Conversation agents fueled by Large Language Models (LLMs) are providing a new way to interact with visual data. While there have been initial attempts for image-based conversation models, this work addresses the under-explored field of…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Muhammad Maaz , Hanoona Rasheed , Salman Khan , Fahad Shahbaz Khan

Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Ziyang Chen , Daniel Geng , Andrew Owens