English
Related papers

Related papers: SAVGBench: Benchmarking Spatially Aligned Audio-Vi…

200 papers

Vision-language models (VLMs) have shown remarkable progress in offline tasks such as image captioning and video question answering. However, real-time interactive environments impose new demands on VLMs, requiring them to generate…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Keunwoo Peter Yu , Joyce Chai

In recent years, artificial intelligence (AI)-driven video generation has gained significant attention. Consequently, there is a growing need for accurate video quality assessment (VQA) metrics to evaluate the perceptual quality of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Zhichao Zhang , Wei Sun , Xinyue Li , Jun Jia , Xiongkuo Min , Zicheng Zhang , Chunyi Li , Zijian Chen , Puyi Wang , Fengyu Sun , Shangling Jui , Guangtao Zhai

Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this work, we challenge this paradigm with WavFlow, a framework that generates high-fidelity…

Audio-Visual Segmentation (AVS) aims to produce pixel-level masks of sound producing objects in videos, by jointly learning from audio and visual signals. However, real-world environments are inherently dynamic, causing audio and visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Siddeshwar Raghavan , Gautham Vinod , Bruce Coburn , Fengqing Zhu

Due to recent advancements in Large Audio-Language Models (LALMs) that demonstrate remarkable performance across a range of sound-, speech- and music-related tasks, there is a growing interest in proposing benchmarks to assess these models.…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-12 Jingru Lin , Chen Zhang , Tianrui Wang , Haizhou Li

With the development of media and networking technologies, multimedia applications ranging from feature presentation in a cinema setting to video on demand to interactive video conferencing are in great demand. Good synchronization between…

Computer Vision and Pattern Recognition · Computer Science 2018-12-17 Naji Khosravan , Shervin Ardeshir , Rohit Puri

Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding objects are often easily recognized based solely on visual…

Multimedia · Computer Science 2024-09-12 Liangyu Chen , Zihao Yue , Boshen Xu , Qin Jin

High-resolution video generation has emerged as a crucial task in computer vision, with wide-ranging applications in entertainment, simulation, and data augmentation. However, generating temporally coherent and visually realistic videos…

Image and Video Processing · Electrical Eng. & Systems 2025-07-08 Abhinav Sagar

Versatile audio super-resolution (SR) aims to predict high-frequency components from low-resolution audio across diverse domains such as speech, music, and sound effects. Existing diffusion-based SR methods often fail to produce…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Jaekwon Im , Juhan Nam

With the rapid advancement of Multi-modal Large Language Models (MLLMs), several diagnostic benchmarks have recently been developed to assess these models' multi-modal reasoning proficiency. However, these benchmarks are restricted to…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Sanjoy Chowdhury , Sayan Nag , Subhrajyoti Dasgupta , Yaoting Wang , Mohamed Elhoseiny , Ruohan Gao , Dinesh Manocha

Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Silin Gao , Sheryl Mathew , Li Mi , Sepideh Mamooler , Mengjie Zhao , Hiromi Wakaki , Yuki Mitsufuji , Syrielle Montariol , Antoine Bosselut

Recent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks. However, the growth is not attributed solely to models and benchmarks. Universally…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Lucas Goncalves , Prashant Mathur , Chandrashekhar Lavania , Metehan Cekic , Marcello Federico , Kyu J. Han

Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video…

Video transitions aim to synthesize intermediate frames between two clips, but naive approaches such as linear blending introduce artifacts that limit professional use or break temporal coherence. Traditional techniques (cross-fades,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Mia Kan , Yilin Liu , Niloy Mitra

This paper presents a baseline approach and an experimental protocol for a specific content verification problem: detecting discrepancies between the audio and video modalities in multimedia content. We first design and optimize an…

Computer Vision and Pattern Recognition · Computer Science 2024-05-02 Konstantinos Apostolidis , Jakob Abesser , Luca Cuccovillo , Vasileios Mezaris

We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioning. The key to our…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Moayed Haji-Ali , Willi Menapace , Aliaksandr Siarohin , Ivan Skorokhodov , Alper Canberk , Kwot Sin Lee , Vicente Ordonez , Sergey Tulyakov

While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over complex narratives remains poorly evaluated. Existing summarization benchmarks fragment…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Mengqi Shi , Haopeng Zhang

The growing capabilities of AI in generating video content have brought forward significant challenges in effectively evaluating these videos. Unlike static images or text, video content involves complex spatial and temporal dynamics which…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Xiao Liu , Xinhao Xiang , Zizhong Li , Yongheng Wang , Zhuoheng Li , Zhuosheng Liu , Weidi Zhang , Weiqi Ye , Jiawei Zhang

Recent video multimodal large language models achieve impressive results across various benchmarks. However, current evaluations suffer from two critical limitations: (1) inflated scores can mask deficiencies in fine-grained visual…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Jiahao Meng , Tan Yue , Qi Xu , Haochen Wang , Zhongwei Ren , Weisong Liu , Yuhao Wang , Renrui Zhang , Yunhai Tong , Haodong Duan

Video and audio content creation serves as the core technique for the movie industry and professional users. Recently, existing diffusion-based methods tackle video and audio generation separately, which hinders the technique transfer from…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Yazhou Xing , Yingqing He , Zeyue Tian , Xintao Wang , Qifeng Chen