English
Related papers

Related papers: AV-Link: Temporally-Aligned Diffusion Features for…

200 papers

Multimodal emotion recognition has recently gained much attention since it can leverage diverse and complementary relationships over multiple modalities (e.g., audio, visual, biosignals, etc.), and can provide some robustness to noisy…

We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Aggelina Chatziagapi , Louis-Philippe Morency , Hongyu Gong , Michael Zollhoefer , Dimitris Samaras , Alexander Richard

This study focuses on a challenging yet promising task, Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text conditions, meanwhile ensuring both modalities are aligned with text. Despite…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Kaisi Guan , Xihua Wang , Zhengfeng Lai , Xin Cheng , Peng Zhang , XiaoJiang Liu , Ruihua Song , Meng Cao

Audio-video generation has often relied on complex multi-stage architectures or sequential synthesis of sound and visuals. We introduce Ovi, a unified paradigm for audio-video generation that models the two modalities as a single generative…

Multimedia · Computer Science 2025-10-03 Chetwin Low , Weimin Wang , Calder Katyal

In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing approaches are constrained by the vocabulary available in their…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Eitan Shaar , Ariel Shaulov , Gal Chechik , Lior Wolf

Video-to-audio (V2A) generation aims to synthesize realistic and semantically aligned audio from silent videos, with potential applications in video editing, Foley sound design, and assistive multimedia. Although the excellent results,…

Video-to-audio (V2A) generation shows great potential in fields such as film production. Despite significant advances, current V2A methods relying on global video information struggle with complex scenes and generating audio tailored to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Yingshan Liang , Keyu Fan , Zhicheng Du , Yiran Wang , Qingyang Shi , Xinyu Zhang , Jiasheng Lu , Peiwu Qin

Despite diffusion models having shown powerful abilities to generate photorealistic images, generating videos that are realistic and diverse still remains in its infancy. One of the key reasons is that current methods intertwine spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Zhiwu Qing , Shiwei Zhang , Jiayu Wang , Xiang Wang , Yujie Wei , Yingya Zhang , Changxin Gao , Nong Sang

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Ziwei Zhou , Zeyuan Lai , Rui Wang , Yifan Yang , Zhen Xing , Yuqing Yang , Qi Dai , Lili Qiu , Chong Luo

Recent audio-video generative systems suggest that coupling modalities benefits not only audio-video synchrony but also the video modality itself. We pose a fundamental question: Does audio-video joint denoising training improve video…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Jianzong Wu , Hao Lian , Dachao Hao , Ye Tian , Qingyu Shi , Biaolong Chen , Hao Jiang , Yunhai Tong

Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing audio-video generation pipelines typically decompose the…

We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from…

Computer Vision and Pattern Recognition · Computer Science 2022-09-30 Uriel Singer , Adam Polyak , Thomas Hayes , Xi Yin , Jie An , Songyang Zhang , Qiyuan Hu , Harry Yang , Oron Ashual , Oran Gafni , Devi Parikh , Sonal Gupta , Yaniv Taigman

Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to three critical requirements: quality of the generated samples, seamless multimodal synchronization and temporal coherence, with audio…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Alex Ergasti , Giuseppe Gabriele Tarollo , Filippo Botti , Tomaso Fontanini , Claudio Ferrari , Massimo Bertozzi , Andrea Prati

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models to learn the…

Sound · Computer Science 2024-03-14 Shentong Mo , Jing Shi , Yapeng Tian

We propose MAViD, a novel Multimodal framework for Audio-Visual Dialogue understanding and generation. Existing approaches primarily focus on non-interactive systems and are limited to producing constrained and unnatural human speech. The…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Youxin Pang , Jiajun Liu , Lingfeng Tan , Yong Zhang , Feng Gao , Xiang Deng , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

In this paper, we present a novel approach to the audio-visual video parsing (AVVP) task that demarcates events from a video separately for audio and visual modalities. The proposed parsing approach simultaneously detects the temporal…

We introduce a novel deep learning-based audio-visual quality (AVQ) prediction model that leverages internal features from state-of-the-art unimodal predictors. Unlike prior approaches that rely on simple fusion strategies, our model…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Ina Salaj , Arijit Biswas

While vision-language pretrained models (VLMs) excel in various multimodal understanding tasks, their potential in fine-grained audio-visual reasoning, particularly for audio-visual question answering (AVQA), remains largely unexplored.…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Yuanyuan Jiang , Jianqin Yin

Sound effect editing-modifying audio by adding, removing, or replacing elements-remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and…

Multimedia · Computer Science 2025-11-27 Xinyue Guo , Xiaoran Yang , Lipan Zhang , Jianxuan Yang , Zhao Wang , Jian Luan

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Hebeizi Li , Zihao Liang , Benyuan Sun , Zihao Yin , Xiao Sha , Chenliang Wang , Yi Yang