中文
相关论文

相关论文: Generating Visually Aligned Sound from Videos

200 篇论文

Modern streaming services are increasingly labeling videos based on their visual or audio content. This typically augments the use of technologies such as AI and ML by allowing to use natural speech for searching by keywords and video…

声音 · 计算机科学 2021-09-22 Ievgeniia Kuzminykh , Dan Shevchuk , Stavros Shiaeles , Bogdan Ghita

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…

The sound effects that designers add to videos are designed to convey a particular artistic effect and, thus, may be quite different from a scene's true sound. Inspired by the challenges of creating a soundtrack for a video that differs…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Yuexi Du , Ziyang Chen , Justin Salamon , Bryan Russell , Andrew Owens

Large multimodal models (LMMs) have shown remarkable progress in audio-visual understanding, yet they struggle with real-world scenarios that require complex reasoning across extensive video collections. Existing benchmarks for video…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Sanjoy Chowdhury , Mohamed Elmoghany , Yohan Abeysinghe , Junjie Fei , Sayan Nag , Salman Khan , Mohamed Elhoseiny , Dinesh Manocha

Neural network-based vocoders have recently demonstrated the powerful ability to synthesize high-quality speech. These models usually generate samples by conditioning on spectral features, such as Mel-spectrogram and fundamental frequency,…

音频与语音处理 · 电气工程与系统科学 2023-03-13 Yunchao He , Yujun Wang

The content of visual and audio scenes is multi-faceted such that a video can be paired with various audio and vice-versa. Thereby, in video-to-audio generation task, it is imperative to introduce steering approaches for controlling the…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Xiulong Liu , Kun Su , Eli Shlizerman

From professional filmmaking to user-generated content, creators and consumers have long recognized that the power of video depends on the harmonious integration of what we hear (the video's audio track) with what we see (the video's image…

Video-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency, and visual-audio temporal synchrony. We propose Frieren, a…

声音 · 计算机科学 2025-01-07 Yongqi Wang , Wenxiang Guo , Rongjie Huang , Jiawei Huang , Zehan Wang , Fuming You , Ruiqi Li , Zhou Zhao

The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings. In this work, we show that ambient sounds can be used as a supervisory signal for learning visual…

计算机视觉与模式识别 · 计算机科学 2017-12-21 Andrew Owens , Jiajun Wu , Josh H. McDermott , William T. Freeman , Antonio Torralba

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to…

计算机视觉与模式识别 · 计算机科学 2022-06-15 Changan Chen , Ruohan Gao , Paul Calamia , Kristen Grauman

Video generation is one of the most challenging tasks in Machine Learning and Computer Vision fields of study. In this paper, we tackle the text to video generation problem, which is a conditional form of video generation. Humans can…

计算机视觉与模式识别 · 计算机科学 2021-07-30 Amir Mazaheri , Mubarak Shah

Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper, we start by…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazim Kemal Ekenel , Alexander Waibel

Audio to Video generation is an interesting problem that has numerous applications across industry verticals including film making, multi-media, marketing, education and others. High-quality video generation with expressive facial movements…

计算机视觉与模式识别 · 计算机科学 2020-12-16 Neeraj Kumar , Srishti Goel , Ankur Narang , Mujtaba Hasan

Video-text retrieval, the task of retrieving videos based on a textual query or vice versa, is of paramount importance for video understanding and multimodal information retrieval. Recent methods in this area rely primarily on visual and…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Boseung Jeong , Jicheol Park , Sungyeon Kim , Suha Kwak

Currently, various studies have been exploring generation of long videos. However, the generated frames in these videos often exhibit jitter and noise. Therefore, in order to generate the videos without these noise, we propose a novel…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Chaoyi Wang , Yaozhe Song , Yafeng Zhang , Jun Pei , Lijie Xia , Jianpo Liu

Generating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently. In this paper, we address this problem with the aid of neural scene representation…

计算机视觉与模式识别 · 计算机科学 2021-08-20 Yudong Guo , Keyu Chen , Sen Liang , Yong-Jin Liu , Hujun Bao , Juyong Zhang

In this paper, we consider a novel and practical case for talking face video generation. Specifically, we focus on the scenarios involving multi-people interactions, where the talking context, such as audience or surroundings, is present.…

计算机视觉与模式识别 · 计算机科学 2024-02-29 Meidai Xuanyuan , Yuwang Wang , Honglei Guo , Qionghai Dai

Automatic transcriptions of consumer-generated multi-media content such as "Youtube" videos still exhibit high word error rates. Such data typically occupies a very broad domain, has been recorded in challenging conditions, with cheap…

计算与语言 · 计算机科学 2017-12-08 Abhinav Gupta , Yajie Miao , Leonardo Neves , Florian Metze

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)…

计算机视觉与模式识别 · 计算机科学 2021-07-21 Suzhen Wang , Lincheng Li , Yu Ding , Changjie Fan , Xin Yu

Temporal consistency is crucial for extending image processing pipelines to the video domain, which is often enforced with flow-based warping error over adjacent frames. Yet for human video synthesis, such scheme is less reliable due to the…

计算机视觉与模式识别 · 计算机科学 2020-12-14 Lingbo Yang , Zhanning Gao , Peiran Ren , Siwei Ma , Wen Gao