中文
相关论文

相关论文: Seeing Voices: Generating A-Roll Video from Audio …

200 篇论文

Lip-to-speech involves generating a natural-sounding speech synchronized with a soundless video of a person talking. Despite recent advances, current methods still cannot produce high-quality speech with high levels of intelligibility for…

音频与语音处理 · 电气工程与系统科学 2024-03-29 Yochai Yemini , Aviv Shamsian , Lior Bracha , Sharon Gannot , Ethan Fetaya

This paper proposes Video-Teller, a video-language foundation model that leverages multi-modal fusion and fine-grained modality alignment to significantly enhance the video-to-text generation task. Video-Teller boosts the training…

计算机视觉与模式识别 · 计算机科学 2023-10-12 Haogeng Liu , Qihang Fan , Tingkai Liu , Linjie Yang , Yunzhe Tao , Huaibo Huang , Ran He , Hongxia Yang

The majority of traditional text-to-video retrieval systems operate in static environments, i.e., there is no interaction between the user and the agent beyond the initial textual query provided by the user. This can be sub-optimal if the…

计算机视觉与模式识别 · 计算机科学 2022-07-19 Avinash Madasu , Junier Oliva , Gedas Bertasius

The increasing amount of online videos brings several opportunities for training self-supervised neural networks. The creation of large scale datasets of videos such as the YouTube-8M allows us to deal with this large amount of data in…

信息检索 · 计算机科学 2018-01-09 Didac Surís , Amanda Duarte , Amaia Salvador , Jordi Torres , Xavier Giró-i-Nieto

Recent video generation models have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and…

计算机视觉与模式识别 · 计算机科学 2026-03-24 David Romero , Ariana Bermudez , Viacheslav Iablochnikov , Hao Li , Fabio Pizzati , Ivan Laptev

Without incurring significant computational overhead, train-free long video generation aims to enable foundation video generation models to produce longer videos. Frame-level autoregressive frameworks, e.g., FIFO-diffusion, offer the…

计算机视觉与模式识别 · 计算机科学 2026-05-19 X. Feng , J. Zhu , M. Wu , C. Chen , F. Mao , H. Guo , J. Wu , X. Chu , K. Huang

Diffusion based video generation has received extensive attention and achieved considerable success within both the academic and industrial communities. However, current efforts are mainly concentrated on single-objective or single-task…

计算机视觉与模式识别 · 计算机科学 2024-01-18 Ludan Ruan , Lei Tian , Chuanwei Huang , Xu Zhang , Xinyan Xiao

Representing wild sounds as images is an important but challenging task due to the lack of paired datasets between sound and images and the significant differences in the characteristics of these two modalities. Previous studies have…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Taegyeong Lee , Jeonghun Kang , Hyeonyu Kim , Taehwan Kim

We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-to-audio models achieve strong semantic…

The increasing use of machine learning models has amplified the demand for high-quality, large-scale multimodal datasets. However, the availability of such datasets, especially those combining acoustic, visual and textual data, remains…

多媒体 · 计算机科学 2025-09-09 Jorge E. León , Miguel Carrasco

It has been shown that learning audiovisual features can lead to improved speech recognition performance over audio-only features, especially for noisy speech. However, in many common applications, the visual features are partially or…

音频与语音处理 · 电气工程与系统科学 2023-12-20 Oscar Chang , Otavio Braga , Hank Liao , Dmitriy Serdyuk , Olivier Siohan

Deep generative models have demonstrated the ability to create realistic audiovisual content, sometimes driven by domains of different nature. However, smooth temporal dynamics in video generation is a challenging problem. This work focuses…

声音 · 计算机科学 2024-06-25 Rafael Redondo

The burgeoning growth of video-to-music generation can be attributed to the ascendancy of multimodal generative models. However, there is a lack of literature that comprehensively combs through the work in this field. To fill this gap, this…

音频与语音处理 · 电气工程与系统科学 2025-12-23 Shulei Ji , Songruoyao Wu , Zihao Wang , Shuyu Li , Kejun Zhang

Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality was approached separately, which poses limitations. The…

While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under-explored frontier. To bridge this gap, we present RISE-Video, a…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Mingxin Liu , Shuran Ma , Shibei Meng , Xiangyu Zhao , Zicheng Zhang , Shaofeng Zhang , Zhihang Zhong , Peixian Chen , Haoyu Cao , Xing Sun , Haodong Duan , Xue Yang

We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabilities such as precise instruction-based video editing and…

Recently, Image-to-Music (I2M) generation has garnered significant attention, with potential applications in fields such as gaming, advertising, and multi-modal art creation. However, due to the ambiguous and subjective nature of I2M tasks,…

声音 · 计算机科学 2026-04-17 Zijian Zhao , Dian Jin , Zijing Zhou

Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally incomplete. Such reliance often yields unnatural or implausible outcomes, especially by…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Lee Hsin-Ying , Hanwen Jiang , Yiqun Mei , Jing Shi , Ming-Hsuan Yang , Zhixin Shu

There have been a number of techniques that have demonstrated the generation of multimedia data for one modality at a time using GANs, such as the ability to generate images, videos, and audio. However, so far, the task of multi-modal…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Vinod K Kurmi , Vipul Bajaj , Badri N Patro , K S Venkatesh , Vinay P Namboodiri , Preethi Jyothi

Prevailing Video-to-Audio (V2A) generation models operate offline, assuming an entire video sequence or chunks of frames are available beforehand. This critically limits their use in interactive applications such as live content creation…