English
Related papers

Related papers: DiffAVA: Personalized Text-to-Audio Generation wit…

200 papers

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

This paper introduces an audio-visual speech enhancement system that leverages score-based generative models, also known as diffusion models, conditioned on visual information. In particular, we exploit audio-visual embeddings obtained from…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-05 Julius Richter , Simone Frintrop , Timo Gerkmann

In this paper, we present JoVA, a unified framework for joint video-audio generation. Despite recent encouraging advances, existing methods face two critical limitations. First, most existing approaches can only generate ambient sounds and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Xiaohu Huang , Hao Zhou , Qiangpeng Yang , Shilei Wen , Kai Han

Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges: (1) the scarcity…

Sound · Computer Science 2026-04-30 Yusheng Dai , Zehua Chen , Yuxuan Jiang , Baolong Gao , Qiuhong Ke , Jianfei Cai , Jun Zhu

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality only. Our idea is to…

Sound · Computer Science 2023-03-15 Changan Chen , Wei Sun , David Harwath , Kristen Grauman

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Zhen Ye , Xu Tan , Aoxiong Yin , Hongzhan Lin , Guangyan Zhang , Peiwen Sun , Yiming Li , Chi-Min Chan , Wei Ye , Shikun Zhang , Wei Xue

A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instantaneous velocity,…

Sound · Computer Science 2025-09-09 Xiaoran Yang , Jianxuan Yang , Xinyue Guo , Haoyu Wang , Ningning Pan , Gongping Huang

Although audio generation shares commonalities across different types of audio, such as speech, music, and sound effects, designing models for each type requires careful consideration of specific objectives and biases that can significantly…

Diffusion models have recently advanced photorealistic human synthesis, although practical talking-head generation (THG) remains constrained by high inference latency, temporal instability such as flicker and identity drift, and imperfect…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Soumya Mazumdar , Vineet Kumar Rakesh

Precisely evaluating semantic alignment between text prompts and generated videos remains a challenge in Text-to-Video (T2V) Generation. Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Kaisi Guan , Zhengfeng Lai , Yuchong Sun , Peng Zhang , Wei Liu , Kieran Liu , Meng Cao , Ruihua Song

In this paper, we introduce a novel task called language-guided joint audio-visual editing. Given an audio and image pair of a sounding event, this task aims at generating new audio-visual content by editing the given sounding event…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Susan Liang , Chao Huang , Yapeng Tian , Anurag Kumar , Chenliang Xu

Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining), have significantly enhanced the alignment of text, video,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Shentong Mo , Zehua Chen , Fan Bao , Jun Zhu

We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently…

Sound · Computer Science 2024-05-27 Xinlei Niu , Jing Zhang , Christian Walder , Charles Patrick Martin

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Yujin Jeong , Yunji Kim , Sanghyuk Chun , Jiyoung Lee

Generating representations of video data is of key importance in advancing the field of machine perception. Most current techniques rely on hand-annotated data, which can be difficult to work with, expensive to generate, and hard to scale.…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Sumanth Gurram , Andy Fang , David Chan , John Canny

The goal of Automatic Voice Over (AVO) is to generate speech in sync with a silent video given its text script. Recent AVO frameworks built upon text-to-speech synthesis (TTS) have shown impressive results. However, the current AVO learning…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-30 Junchen Lu , Berrak Sisman , Mingyang Zhang , Haizhou Li

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Jaemin Jung , Junseok Ahn , Chaeyoung Jung , Tan Dat Nguyen , Youngjoon Jang , Joon Son Chung

Diffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited…

Sound · Computer Science 2023-08-04 Ke Chen , Yusong Wu , Haohe Liu , Marianna Nezhurina , Taylor Berg-Kirkpatrick , Shlomo Dubnov

Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However, video-to-audio dubbing for long-form content remains an…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yehang Zhang , Xinli Xu , Xiaojie Xu , Li Liu , Yingcong Chen

Zero-shot learning enables models to generalise to unseen classes by leveraging semantic information, bridging the gap between training and testing sets with non-overlapping classes. While much research has focused on zero-shot learning in…

Sound · Computer Science 2025-07-03 Ysobel Sims , Alexandre Mendes , Stephan Chalup
‹ Prev 1 3 4 5 6 7 10 Next ›