English
Related papers

Related papers: Foley-Flow: Coordinated Video-to-Audio Generation …

200 papers

Recent advancements in discrete token-based speech generation have highlighted the importance of token-to-waveform generation for audio quality, particularly in real-time interactions. Traditional frameworks integrating semantic tokens with…

Sound · Computer Science 2025-07-02 Dake Guo , Jixun Yao , Linhan Ma , He Wang , Lei Xie

Current visual generation methods can produce high quality videos guided by texts. However, effectively controlling object dynamics remains a challenge. This work explores audio as a cue to generate temporally synchronized image animations.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Lin Zhang , Shentong Mo , Yijing Zhang , Pedro Morgado

Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape. Despite advancements in AIGC technologies for text and image generation, the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-18 Ruibo Fu , Shuchen Shi , Hongming Guo , Tao Wang , Chunyu Qiang , Zhengqi Wen , Jianhua Tao , Xin Qi , Yi Lu , Xiaopeng Wang , Zhiyong Wang , Yukun Liu , Xuefei Liu , Shuai Zhang , Guanjun Li

Visual and acoustic events in the physical world are inherently coupled, yet existing video editing methods typically adopt decoupled pipelines, lacking bidirectional modality interaction. This results in two key limitations: (i)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Sen Liang , Cong Wang , Fengbin Guan , Zhentao Yu , Yiting Lu , Yuanzhi Wang , Yuan Zhou , Xin Li , Zhibo Chen

Echocardiography is widely used for assessing cardiac function, where clinically meaningful parameters such as left-ventricular ejection fraction (EF) play a central role in diagnosis and management. Generative models capable of…

Image and Video Processing · Electrical Eng. & Systems 2026-03-17 Emmanuel Oladokun , Sarina Thomas , Jurica Šprem , Vicente Grau

Video-to-Audio (V2A) Generation achieves significant progress and plays a crucial role in film and video post-production. However, current methods overlook the cinematic language, a critical component of artistic expression in filmmaking.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Feizhen Huang , Yu Wu , Yutian Lin , Bo Du

Video to sound generation aims to generate realistic and natural sound given a video input. However, previous video-to-sound generation methods can only generate a random or average timbre without any controls or specializations of the…

Multimedia · Computer Science 2022-11-22 Chenye Cui , Yi Ren , Jinglin Liu , Rongjie Huang , Zhou Zhao

Existing dominant methods for audio generation include Generative Adversarial Networks (GANs) and diffusion-based methods like Flow Matching. GANs suffer from slow convergence during training, while diffusion methods require multi-step…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Zengwei Yao , Wei Kang , Han Zhu , Liyong Guo , Lingxuan Ye , Fangjun Kuang , Weiji Zhuang , Zhaoqing Li , Zhifeng Han , Long Lin , Daniel Povey

Video-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency, and visual-audio temporal synchrony. We propose Frieren, a…

Sound · Computer Science 2025-01-07 Yongqi Wang , Wenxiang Guo , Rongjie Huang , Jiawei Huang , Zehan Wang , Fuming You , Ruiqi Li , Zhou Zhao

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced audio codecs, the…

LLM-conditioned segmentation has recently advanced rapidly by coupling large language models with iterative mask generation frameworks. However, we identify a persistent failure mode in current propose-then-select pipelines. Although…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Zekang Zhang , Guangyu Gao , Youyun Tang , ChengJing Wu , Xiaochao Qu , Chi Harold Liu , Jianbo Jiao , Yunchao Wei , Luoqi Liu , Ting Liu

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Zhe Cao , Tao Wang , Jiaming Wang , Yanghai Wang , Yuanxing Zhang , Jialu Chen , Miao Deng , Jiahao Wang , Yubin Guo , Chenxi Liao , Yize Zhang , Zhaoxiang Zhang , Jiaheng Liu

We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on visual generative…

Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge this critical gap, we present an entropy-aware gated fusion…

Multimedia · Computer Science 2025-05-29 Le Xu , Chenxing Li , Yong Ren , Yujie Chen , Yu Gu , Ruibo Fu , Shan Yang , Dong Yu

We introduce a multi-modal diffusion model tailored for the bi-directional conditional generation of video and audio. We propose a joint contrastive training loss to improve the synchronization between visual and auditory occurrences. We…

Machine Learning · Computer Science 2024-10-10 Ruihan Yang , Hannes Gamper , Sebastian Braun

Cross-modal audio-visual perception has been a long-lasting topic in psychology and neurology, and various studies have discovered strong correlations in human perception of auditory and visual stimuli. Despite works in computational…

Computer Vision and Pattern Recognition · Computer Science 2017-04-28 Lele Chen , Sudhanshu Srivastava , Zhiyao Duan , Chenliang Xu

People can easily imagine the potential sound while seeing an event. This natural synchronization between audio and visual signals reveals their intrinsic correlations. To this end, we propose to learn the audio-visual correlations from the…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Ye Zhu , Yu Wu , Hugo Latapie , Yi Yang , Yan Yan

In this paper, we abstract the process of people hearing speech, extracting meaningful cues, and creating various dynamically audio-consistent talking faces, termed Listening and Imagining, into the task of high-fidelity diverse talking…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Chao Xu , Yang Liu , Jiazheng Xing , Weida Wang , Mingze Sun , Jun Dan , Tianxin Huang , Siyuan Li , Zhi-Qi Cheng , Ying Tai , Baigui Sun

Audio-Visual Segmentation (AVS) aims to precisely outline audible objects in a visual scene at the pixel level. Existing AVS methods require fine-grained annotations of audio-mask pairs in supervised learning fashion. This limits their…

Computer Vision and Pattern Recognition · Computer Science 2023-09-14 Swapnil Bhosale , Haosen Yang , Diptesh Kanojia , Xiatian Zhu

Audio is an essential part of our life, but creating it often requires expertise and is time-consuming. Research communities have made great progress over the past year advancing the performance of large scale audio generative models for a…

‹ Prev 1 4 5 6 7 8 10 Next ›