English
Related papers

Related papers: Video Background Music Generation: Dataset, Method…

200 papers

The next frontier for video generation lies in developing models capable of zero-shot reasoning, where understanding real-world scientific laws is crucial for accurate physical outcome modeling under diverse conditions. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Lanxiang Hu , Abhilash Shankarampeta , Yixin Huang , Zilin Dai , Haoyang Yu , Yujie Zhao , Haoqiang Kang , Daniel Zhao , Tajana Rosing , Hao Zhang

Machine-generated music (MGM) has emerged as a powerful tool with applications in music therapy, personalised editing, and creative inspiration for the music community. However, its unregulated use threatens the entertainment, education,…

Sound · Computer Science 2026-02-16 Yupei Li , Hanqian Li , Lucia Specia , Björn W. Schuller

Video generation has many unique challenges beyond those of image generation. The temporal dimension introduces extensive possible variations across frames, over which consistency and continuity may be violated. In this study, we move…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Weixi Feng , Jiachen Li , Michael Saxon , Tsu-jui Fu , Wenhu Chen , William Yang Wang

Video-quality measurement is a critical task in video processing. Nowadays, many implementations of new encoding standards - such as AV1, VVC, and LCEVC - use deep-learning-based decoding algorithms with perceptual metrics that serve as…

Computer Vision and Pattern Recognition · Computer Science 2023-02-08 Anastasia Antsiferova , Sergey Lavrushkin , Maksim Smirnov , Alexander Gushchin , Dmitriy Vatolin , Dmitriy Kulikov

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

Massive multi-modality datasets play a significant role in facilitating the success of large video-language models. However, current video-language datasets primarily provide text descriptions for visual frames, considering audio to be…

The vision and language generative models have been overgrown in recent years. For video generation, various open-sourced models and public-available services have been developed to generate high-quality videos. However, these methods often…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Yaofang Liu , Xiaodong Cun , Xuebo Liu , Xintao Wang , Yong Zhang , Haoxin Chen , Yang Liu , Tieyong Zeng , Raymond Chan , Ying Shan

In this work, we study music/video cross-modal recommendation, i.e. recommending a music track for a video or vice versa. We rely on a self-supervised learning paradigm to learn from a large amount of unlabelled data. We rely on a…

Multimedia · Computer Science 2021-05-03 Laure Pretet , Gael Richard , Geoffroy Peeters

Recent years have seen many audio-domain text-to-music generation models that rely on large amounts of text-audio pairs for training. However, symbolic-domain controllable music generation has lagged behind partly due to the lack of a…

Sound · Computer Science 2025-06-17 Weihan Xu , Julian McAuley , Taylor Berg-Kirkpatrick , Shlomo Dubnov , Hao-Wen Dong

Content creators often use music to enhance their videos, from soundtracks in movies to background music in video blogs and social media content. However, identifying the best music for a video can be a difficult and time-consuming task. To…

Multimedia · Computer Science 2024-12-24 Shanti Stewart , Gouthaman KV , Lie Lu , Andrea Fanelli

Video description is the automatic generation of natural language sentences that describe the contents of a given video. It has applications in human-robot interaction, helping the visually impaired and video subtitling. The past few years…

Computer Vision and Pattern Recognition · Computer Science 2020-03-04 Nayyer Aafaq , Ajmal Mian , Wei Liu , Syed Zulqarnain Gilani , Mubarak Shah

Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can…

We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing benchmarks focused on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Dahun Kim , AJ Piergiovanni , Ganesh Mallya , Anelia Angelova

LED Virtual Production (VP) uses large LED volumes to render backgrounds in real time, enabling in-camera visual effects but making post-shot changes labor-intensive. We address this with CineMatte, a robust background matting framework for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yuanjian He , Chen Zhang , Fasheng Chen , Jiangbo Cao

Recently, the witness of the rapidly growing popularity of short videos on different Internet platforms has intensified the need for a background music (BGM) retrieval system. However, existing video-music retrieval methods only based on…

Information Retrieval · Computer Science 2021-08-04 Tingtian Li , Zixun Sun , Haoruo Zhang , Jin Li , Ziming Wu , Hui Zhan , Yipeng Yu , Hengcan Shi

In 2021, a new track has been initiated in the Challenge for Learned Image Compression~: the video track. This category proposes to explore technologies for the compression of short video clips at 1 Mbit/s. This paper proposes to generate…

Image and Video Processing · Electrical Eng. & Systems 2021-05-21 Théo Ladune , Pierrick Philippe

Most existing video diffusion models (VDMs) are limited to mere text conditions. Thereby, they are usually lacking in control over visual appearance and geometry structure of the generated videos. This work presents Moonshot, a new video…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 David Junhao Zhang , Dongxu Li , Hung Le , Mike Zheng Shou , Caiming Xiong , Doyen Sahoo

As research on neural volumetric video reconstruction and compression flourishes, there is a need for diverse and realistic datasets, which can be used to develop and validate reconstruction and compression models. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Adrian Azzarelli , Ge Gao , Ho Man Kwan , Fan Zhang , Nantheera Anantrasirichai , Ollie Moolan-Feroze , David Bull

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Ziqi Huang , Yinan He , Jiashuo Yu , Fan Zhang , Chenyang Si , Yuming Jiang , Yuanhan Zhang , Tianxing Wu , Qingyang Jin , Nattapol Chanpaisit , Yaohui Wang , Xinyuan Chen , Limin Wang , Dahua Lin , Yu Qiao , Ziwei Liu

We describe the implementation of and early results from a system that automatically composes picture-synched musical soundtracks for videos and movies. We use the phrase "picture-synched" to mean that the structure of the automatically…

Multimedia · Computer Science 2019-10-22 Vansh Dassani , Jon Bird , Dave Cliff