English
Related papers

Related papers: Generating Visually Aligned Sound from Videos

200 papers

Temporal video alignment aims to synchronize the key events like object interactions or action phase transitions in two videos. Such methods could benefit various video editing, processing, and understanding tasks. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Ishan Rajendrakumar Dave , Fabian Caba Heilbron , Mubarak Shah , Simon Jenni

All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise to the utterance or…

Computer Vision and Pattern Recognition · Computer Science 2019-10-03 Gaurav Mittal , Baoyuan Wang

Speechreading is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible acoustic speech signal from silent video frames…

Computer Vision and Pattern Recognition · Computer Science 2017-01-10 Ariel Ephrat , Shmuel Peleg

Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal…

Sound · Computer Science 2024-09-16 Zhiqi Huang , Dan Luo , Jun Wang , Huan Liao , Zhiheng Li , Zhiyong Wu

The advancement of generation models has led to the emergence of highly realistic artificial intelligence (AI)-generated videos. Malicious users can easily create non-existent videos to spread false information. This letter proposes an…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Jianfa Bai , Man Lin , Gang Cao

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of convolutional…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Sudha Krishnamurthy

Perceiving meaningful activities in a long video sequence is a challenging problem due to ambiguous definition of 'meaningfulness' as well as clutters in the scene. We approach this problem by learning a generative model for regular motion…

Computer Vision and Pattern Recognition · Computer Science 2016-04-18 Mahmudul Hasan , Jonghyun Choi , Jan Neumann , Amit K. Roy-Chowdhury , Larry S. Davis

Despite recent advances in retrieval-augmented generation (RAG) for video understanding, effectively understanding long-form video content remains underexplored due to the vast scale and high complexity of video data. Current RAG approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Nianbo Zeng , Haowen Hou , Fei Richard Yu , Si Shi , Ying Tiffany He

Video diffusion models are able to generate high-quality videos by learning strong spatial-temporal priors on large-scale datasets. In this paper, we aim to investigate whether such priors derived from a generative process are suitable for…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Zejia Weng , Xitong Yang , Zhen Xing , Zuxuan Wu , Yu-Gang Jiang

Modern video generators still struggle with complex physical dynamics, often falling short of physical realism. Existing approaches address this using external verifiers or additional training on augmented data, which is computationally…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Sangwon Jang , Taekyung Ki , Jaehyeong Jo , Saining Xie , Jaehong Yoon , Sung Ju Hwang

Despite advancements in artificial intelligence, object recognition models still lag behind in emulating visual information processing in human brains. Recent studies have highlighted the potential of using neural data to mimic brain…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Zitong Lu , Yile Wang , Julie D. Golomb

The objective of this work is to localize sound sources that are visible in a video without using manual annotations. Our key technical contribution is to show that, by training the network to explicitly discriminate challenging image…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Honglie Chen , Weidi Xie , Triantafyllos Afouras , Arsha Nagrani , Andrea Vedaldi , Andrew Zisserman

Deepfakes is a branch of malicious techniques that transplant a target face to the original one in videos, resulting in serious problems such as infringement of copyright, confusion of information, or even public panic. Previous efforts for…

Computer Vision and Pattern Recognition · Computer Science 2021-04-12 Zekun Sun , Yujie Han , Zeyu Hua , Na Ruan , Weijia Jia

Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelity lip-sync with different techniques, their task-orientated…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Jiazhi Guan , Zhiliang Xu , Hang Zhou , Kaisiyuan Wang , Shengyi He , Zhanwang Zhang , Borong Liang , Haocheng Feng , Errui Ding , Jingtuo Liu , Jingdong Wang , Youjian Zhao , Ziwei Liu

We introduce VampNet, a masked acoustic token modeling approach to music synthesis, compression, inpainting, and variation. We use a variable masking schedule during training which allows us to sample coherent music from the model by…

Sound · Computer Science 2023-07-13 Hugo Flores Garcia , Prem Seetharaman , Rithesh Kumar , Bryan Pardo

Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual…

Sound · Computer Science 2024-10-18 Ruiqi Li , Siqi Zheng , Xize Cheng , Ziang Zhang , Shengpeng Ji , Zhou Zhao

In this paper, we present MovieFactory, a powerful framework to generate cinematic-picture (3072$\times$1280), film-style (multi-scene), and multi-modality (sounding) movies on the demand of natural languages. As the first fully automated…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Junchen Zhu , Huan Yang , Huiguo He , Wenjing Wang , Zixi Tuo , Wen-Huang Cheng , Lianli Gao , Jingkuan Song , Jianlong Fu

We introduce the novel problem of automatically generating animated GIFs from video. GIFs are short looping video with no sound, and a perfect combination between image and video that really capture our attention. GIFs tell a story, express…

Computer Vision and Pattern Recognition · Computer Science 2016-05-17 Michael Gygli , Yale Song , Liangliang Cao

This paper presents a simple method for speech videos generation based on audio: given a piece of audio, we can generate a video of the target face speaking this audio. We propose Generative Adversarial Networks (GAN) with cut speech audio…

Sound · Computer Science 2022-07-20 Hanhaodi Zhang

Human speech is often accompanied by body gestures including arm and hand gestures. We present a method that reenacts a high-quality video with gestures matching a target speech audio. The key idea of our method is to split and re-assemble…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Yang Zhou , Jimei Yang , Dingzeyu Li , Jun Saito , Deepali Aneja , Evangelos Kalogerakis