中文
相关论文

相关论文: MultiSoundGen: Video-to-Audio Generation for Multi…

200 篇论文

This paper introduces V2A-DPO, a novel Direct Preference Optimization (DPO) framework tailored for flow-based video-to-audio generation (V2A) models, incorporating key adaptations to effectively align generated audio with human preferences.…

声音 · 计算机科学 2026-03-13 Nolan Chan , Timmy Gang , Yongqian Wang , Yuzhe Liang , Dingdong Wang

Video-to-audio (V2A) generation aims to produce corresponding audio given silent video inputs. This task is particularly challenging due to the cross-modality and sequential nature of the audio-visual features involved. Recent works have…

声音 · 计算机科学 2024-09-17 Mingjing Yi , Ming Li

Recent advancements in video-audio joint generation have achieved remarkable success in semantic correspondence. However, achieving precise temporal synchronization, which requires fine-grained alignment between audio events and their…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Xin Cheng , Xihua Wang , Ying Ba , Yuyue Wang , Kaisi Guan , Yinbo Wang , Wenpu Li , Ruihua Song

The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have limited generation…

声音 · 计算机科学 2023-07-03 Simian Luo , Chuanhao Yan , Chenxu Hu , Hang Zhao

Currently, high-quality, synchronized audio is synthesized using various multi-modal joint learning frameworks, leveraging video and optional text inputs. In the video-to-audio benchmarks, video-to-audio quality, semantic alignment, and…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Haomin Zhang , Chang Liu , Junjie Zheng , Zihao Chen , Chaofan Ding , Xinhan Di

With the rapid development of AIGC technology, significant progress has been made in diffusion model-based technologies for text-to-image (T2I) and text-to-video (T2V). In recent years, a few studies have introduced the strategy of Direct…

计算机视觉与模式识别 · 计算机科学 2025-02-05 Lifan Jiang , Boxi Wu , Jiahui Zhang , Xiaotong Guan , Shuang Chen

Multimodality-to-Multiaudio (MM2MA) generation faces significant challenges in synthesizing diverse and contextually aligned audio types (e.g., sound effects, speech, music, and songs) from multimodal inputs (e.g., video, text, images),…

声音 · 计算机科学 2025-08-06 Yan Rong , Jinting Wang , Guangzhi Lei , Shan Yang , Li Liu

Recent advancements in audio generation have been spurred by the evolution of large-scale deep learning models and expansive datasets. However, the task of video-to-audio (V2A) generation continues to be a challenge, principally because of…

音频与语音处理 · 电气工程与系统科学 2023-09-20 Xinhao Mei , Varun Nagaraja , Gael Le Lan , Zhaoheng Ni , Ernie Chang , Yangyang Shi , Vikas Chandra

Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers significant application flexibility, yet faces two unexplored foundational challenges: (1) the scarcity…

声音 · 计算机科学 2026-04-30 Yusheng Dai , Zehua Chen , Yuxuan Jiang , Baolong Gao , Qiuhong Ke , Jianfei Cai , Jun Zhu

While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This paper proposes a new task, video-to-soundtrack (V2ST)…

多媒体 · 计算机科学 2025-07-15 Wenjie Tian , Xinfa Zhu , Haohe Liu , Zhixian Zhao , Zihao Chen , Chaofan Ding , Xinhan Di , Junjie Zheng , Lei Xie

Video-to-Audio (V2A) generation is essential for immersive multimedia experiences, yet its evaluation remains underexplored. Existing benchmarks typically assess diverse audio types under a unified protocol, overlooking the fine-grained…

声音 · 计算机科学 2026-04-14 Qian Zhang , Yuqin Cao , Yixuan Gao , Xiongkuo Min

Prevailing Video-to-Audio (V2A) generation models operate offline, assuming an entire video sequence or chunks of frames are available beforehand. This critically limits their use in interactive applications such as live content creation…

Video-to-audio (V2A) generation is important for video editing and post-processing, enabling the creation of semantics-aligned audio for silent video. However, most existing methods focus on generating short-form audio for short video…

声音 · 计算机科学 2024-12-31 Xin Cheng , Xihua Wang , Yihan Wu , Yuyue Wang , Ruihua Song

Recent Video-to-Audio (V2A) generation relies on extracting semantic and temporal features from video to condition generative models. Training these models from scratch is resource intensive. Consequently, leveraging foundation models (FMs)…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Gehui Chen , Guan'an Wang , Xiaowen Huang , Jitao Sang

This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Junwon Lee , Juhan Nam , Jiyoung Lee

Video-to-audio (V2A) generation aims to synthesize realistic and semantically aligned audio from silent videos, with potential applications in video editing, Foley sound design, and assistive multimedia. Although the excellent results,…

Video-to-Audio (V2A) generation requires balancing four critical perceptual dimensions: semantic consistency, audio-visual temporal synchrony, aesthetic quality, and spatial accuracy; yet existing methods suffer from objective entanglement…

声音 · 计算机科学 2026-03-04 Huadai Liu , Kaicheng Luo , Wen Wang , Qian Chen , Peiwen Sun , Rongjie Huang , Xiangang Li , Jieping Ye , Wei Xue

AIGC has rapidly expanded from text-to-image generation toward high-quality multimodal synthesis across video and audio. Within this context, joint audio-video generation (JAVG) has emerged as a fundamental task that produces synchronized…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Kai Liu , Yanhao Zheng , Kai Wang , Shengqiong Wu , Rongjunchen Zhang , Jiebo Luo , Dimitrios Hatzinakos , Ziwei Liu , Hao Fei , Tat-Seng Chua

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader…

声音 · 计算机科学 2025-03-25 Yong Ren , Chenxing Li , Manjie Xu , Wei Liang , Yu Gu , Rilin Chen , Dong Yu

Building artificial intelligence (AI) systems on top of a set of foundation models (FMs) is becoming a new paradigm in AI research. Their representative and generative abilities learnt from vast amounts of data can be easily adapted and…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Heng Wang , Jianbo Ma , Santiago Pascual , Richard Cartwright , Weidong Cai
‹ 上一页 1 2 3 10 下一页 ›