English
Related papers

Related papers: FoleyGAN: Visually Guided Generative Adversarial N…

200 papers

Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous speakers across…

Sound · Computer Science 2023-02-17 Sang-gil Lee , Wei Ping , Boris Ginsburg , Bryan Catanzaro , Sungroh Yoon

We propose a conditional generative adversarial network (GAN) model for zero-shot video generation. In this study, we have explored zero-shot conditional generation setting. In other words, we generate unseen videos from training samples…

Computer Vision and Pattern Recognition · Computer Science 2021-09-14 Shun Kimura , Kazuhiko Kawamoto

Synthetic creation of drum sounds (e.g., in drum machines) is commonly performed using analog or digital synthesis, allowing a musician to sculpt the desired timbre modifying various parameters. Typically, such parameters control low-level…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-29 J. Nistal , S. Lattner , G. Richard

The research in Environmental Sound Classification (ESC) has been progressively growing with the emergence of deep learning algorithms. However, data scarcity poses a major hurdle for any huge advance in this domain. Data augmentation…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-16 Aswathy Madhu , Suresh K

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far…

Sound · Computer Science 2025-03-31 Yunming Liang , Zihao Chen , Chaofan Ding , Xinhan Di

Since the introduction of Generative Adversarial Networks (GANs) [Goodfellow et al., 2014] there has been a regular stream of both technical advances (e.g., Arjovsky et al. [2017]) and creative uses of these generative models (e.g., [Karras…

Sound · Computer Science 2020-11-11 Pablo Samuel Castro

In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Chenxu Zhang , Yifan Zhao , Yifei Huang , Ming Zeng , Saifeng Ni , Madhukar Budagavi , Xiaohu Guo

We tackle the problem of generating audio samples conditioned on descriptive text captions. In this work, we propose AaudioGen, an auto-regressive generative model that generates audio samples conditioned on text inputs. AudioGen operates…

Time series synthesis is an effective approach to ensuring the secure circulation of time series data. Existing time series synthesis methods typically perform temporal modeling based on random sequences to generate target sequences, which…

Machine Learning · Computer Science 2025-09-01 Xuan Hou , Shuhan Liu , Zhaohui Peng , Yaohui Chu , Yue Zhang , Yining Wang

Deep learning models have demonstrated high-quality performance in areas such as image classification and speech processing. However, creating a deep learning model using electronic health record (EHR) data, requires addressing particular…

Machine Learning · Computer Science 2020-03-06 Amirsina Torfi , Edward A. Fox

Conditional generative adversarial networks (cGAN) have led to large improvements in the task of conditional image generation, which lies at the heart of computer vision. The major focus so far has been on performance improvement, while…

Machine Learning · Computer Science 2019-03-14 Grigorios G. Chrysos , Jean Kossaifi , Stefanos Zafeiriou

The synthesis of synchronized audio-visual content is a key challenge in generative AI, with open-source models facing challenges in robust audio-video alignment. Our analysis reveals that this issue is rooted in three fundamental…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Teng Hu , Zhentao Yu , Guozhen Zhang , Zihan Su , Zhengguang Zhou , Youliang Zhang , Yuan Zhou , Qinglin Lu , Ran Yi

Adversarially trained generative models (GANs) have recently achieved compelling image synthesis results. But despite early successes in using GANs for unsupervised representation learning, they have since been superseded by approaches…

Computer Vision and Pattern Recognition · Computer Science 2019-11-06 Jeff Donahue , Karen Simonyan

Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley task, presents…

Sound · Computer Science 2024-09-12 Qi Yang , Binjie Mao , Zili Wang , Xing Nie , Pengfei Gao , Ying Guo , Cheng Zhen , Pengfei Yan , Shiming Xiang

We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce the target video, then…

Multimedia · Computer Science 2026-03-18 Masato Ishii , Akio Hayakawa , Takashi Shibuya , Yuki Mitsufuji

Speech is a rich biometric signal that contains information about the identity, gender and emotional state of the speaker. In this work, we explore its potential to generate face images of a speaker by conditioning a Generative Adversarial…

Improving speech system performance in noisy environments remains a challenging task, and speech enhancement (SE) is one of the effective techniques to solve the problem. Motivated by the promising results of generative adversarial networks…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-05 Daniel Michelsanti , Zheng-Hua Tan

Generative adversarial networks (GAN) present state-of-the-art results in the generation of samples following the distribution of the input dataset. However, GANs are difficult to train, and several aspects of the model should be previously…

Neural and Evolutionary Computing · Computer Science 2019-12-16 Victor Costa , Nuno Lourenço , João Correia , Penousal Machado

Video generation is a challenging task that requires modeling plausible spatial and temporal dynamics in a video. Inspired by how humans perceive a video by grouping a scene into moving and stationary components, we propose a method that…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Arti Keshari , Sonam Gupta , Sukhendu Das

Can we develop a model that can synthesize realistic speech directly from a latent space, without explicit conditioning? Despite several efforts over the last decade, previous adversarial and diffusion-based approaches still struggle to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-26 Matthew Baas , Herman Kamper
‹ Prev 1 3 4 5 6 7 10 Next ›