中文
相关论文

相关论文: SALSA-V: Shortcut-Augmented Long-form Synchronized…

200 篇论文

Sounding Video Generation (SVG) is an audio-video joint generation task challenged by high-dimensional signal spaces, distinct data formats, and different patterns of content information. To address these issues, we introduce a novel…

计算机视觉与模式识别 · 计算机科学 2024-10-03 Mingzhen Sun , Weining Wang , Yanyuan Qiao , Jiahui Sun , Zihan Qin , Longteng Guo , Xinxin Zhu , Jing Liu

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end av-LLM for video…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Guangzhi Sun , Wenyi Yu , Changli Tang , Xianzhao Chen , Tian Tan , Wei Li , Lu Lu , Zejun Ma , Yuxuan Wang , Chao Zhang

While diffusion model for audio-driven avatar video generation have achieved notable process in synthesizing long sequences with natural audio-visual synchronization and identity consistency, the generation of music-performance videos with…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Jiahui Chen , Weida Wang , Runhua Shi , Huan Yang , Chaofan Ding , Zihao Chen

This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Junwon Lee , Juhan Nam , Jiyoung Lee

Video-to-audio (V2A) generation aims to synthesize realistic and semantically aligned audio from silent videos, with potential applications in video editing, Foley sound design, and assistive multimedia. Although the excellent results,…

Latent diffusion models have shown promising results in audio generation, making notable advancements over traditional methods. However, their performance, while impressive with short audio clips, faces challenges when extended to longer…

声音 · 计算机科学 2024-07-16 Zhenxiong Tan , Xinyin Ma , Gongfan Fang , Xinchao Wang

Speech enhancement systems are typically trained using pairs of clean and noisy speech. In audio-visual speech enhancement (AVSE), there is not as much ground-truth clean data available; most audio-visual datasets are collected in…

音频与语音处理 · 电气工程与系统科学 2024-11-05 Ju-Chieh Chou , Chung-Ming Chien , Karen Livescu

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap…

音频与语音处理 · 电气工程与系统科学 2025-03-24 Ji-Hoon Kim , Jeongsoo Choi , Jaehun Kim , Chaeyoung Jung , Joon Son Chung

Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specific training videos,…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Lin Zhang , Zefan Cai , Yufan Zhou , Shentong Mo , Jinhong Lin , Cheng-En Wu , Yibing Wei , Yijing Zhang , Ruiyi Zhang , Wen Xiao , Tong Sun , Junjie Hu , Pedro Morgado

Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Juhyeong Seon , Woobin Im , Sebin Lee , Jumin Lee , Sung-Eui Yoon

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

The primary aim of Audio-Visual Segmentation (AVS) is to precisely identify and locate auditory elements within visual scenes by accurately predicting segmentation masks at the pixel level. Achieving this involves comprehensively…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Khanh-Binh Nguyen , Chae Jung Park

Recent advances in video generation have achieved remarkable improvements in visual content fidelity. However, the absence of synchronized audio severely undermines immersive experience and restricts practical applications of these…

声音 · 计算机科学 2025-12-09 Fu Li , Weichao Zhao , You Li , Zhichao Zhou , Dongliang He

Audio to Video generation is an interesting problem that has numerous applications across industry verticals including film making, multi-media, marketing, education and others. High-quality video generation with expressive facial movements…

计算机视觉与模式识别 · 计算机科学 2020-12-16 Neeraj Kumar , Srishti Goel , Ankur Narang , Mujtaba Hasan

Sound design involves creatively selecting, recording, and editing sound effects for various media like cinema, video games, and virtual/augmented reality. One of the most time-consuming steps when designing sound is synchronizing audio…

We present xGen-VideoSyn-1, a text-to-video (T2V) generation model capable of producing realistic scenes from textual descriptions. Building on recent advancements, such as OpenAI's Sora, we explore the latent diffusion model (LDM)…

Versatile audio super-resolution (SR) aims to predict high-frequency components from low-resolution audio across diverse domains such as speech, music, and sound effects. Existing diffusion-based SR methods often fail to produce…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Jaekwon Im , Juhan Nam

Endeavors have been made to explore Large Language Models for video analysis (Video-LLMs), particularly in understanding and interpreting long videos. However, existing Video-LLMs still face challenges in effectively integrating the rich…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Jungang Li , Sicheng Tao , Yibo Yan , Xiaojie Gu , Haodong Xu , Xu Zheng , Yuanhuiyi Lyu , Linfeng Zhang , Xuming Hu

Sounding Video Generation (SVG) remains a challenging task due to the inherent structural misalignment between audio and video, as well as the high computational cost of multimodal data processing. In this paper, we introduce ProAV-DiT, a…

多媒体 · 计算机科学 2025-11-18 Jiahui Sun , Weining Wang , Mingzhen Sun , Yirong Yang , Xinxin Zhu , Jing Liu

Generating long, high-quality videos remains a challenge due to the complex interplay of spatial and temporal dynamics and hardware limitations. In this work, we introduce MaskFlow, a unified video generation framework that combines…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Michael Fuest , Vincent Tao Hu , Björn Ommer