中文
相关论文

相关论文: Secure & Personalized Music-to-Video Generation vi…

200 篇论文

We propose an end-to-end pipeline for both building and tracking 3D facial models from personalized in-the-wild (cellphone, webcam, youtube clips, etc.) video data. First, we present a method for automatic data curation and retrieval based…

计算机视觉与模式识别 · 计算机科学 2022-04-28 Winnie Lin , Yilin Zhu , Demi Guo , Ron Fedkiw

With the growing demand for short videos and personalized content, automated Video Log (Vlog) generation has become a key direction in multimodal content creation. Existing methods mostly rely on predefined scripts, lacking dynamism and…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Xiaolu Hou , Bing Ma , Jiaxiang Cheng , Xuhua Ren , Kai Yu , Wenyue Li , Tianxiang Zheng , Qinglin Lu

In this work, we propose a complete framework that generates visual art. Unlike previous stylization methods that are not flexible with style parameters (i.e., they allow stylization with only one style image, a single stylization text or…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Marian Lupascu , Ryan Murdock , Ionut Mironica , Yijun Li

Fast and user-controllable music generation could enable novel ways of composing or performing music. However, state-of-the-art music generation systems require large amounts of data and computational resources for training, and are slow at…

声音 · 计算机科学 2022-08-19 Marco Pasini , Jan Schlüter

In an era where visual content generation is increasingly driven by machine learning, the integration of human feedback into generative models presents significant opportunities for enhancing user experience and output quality. This study…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Dimitri von Rütte , Elisabetta Fedele , Jonathan Thomm , Lukas Wolf

Generating realistic talking faces is an interesting and long-standing topic in the field of computer vision. Although significant progress has been made, it is still challenging to generate high-quality dynamic faces with personalized…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Bo Ding , Zhenfeng Fan , Shuang Yang , Shihong Xia

We investigate how to enhance the physical fidelity of video generation models by leveraging synthetic videos derived from computer graphics pipelines. These rendered videos respect real-world physics, such as maintaining 3D consistency,…

图像与视频处理 · 电气工程与系统科学 2025-03-28 Qi Zhao , Xingyu Ni , Ziyu Wang , Feng Cheng , Ziyan Yang , Lu Jiang , Bohan Wang

Recognizing pain in video is crucial for improving patient-computer interaction systems, yet traditional data collection in this domain raises significant ethical and logistical challenges. This study introduces a novel approach that…

计算机视觉与模式识别 · 计算机科学 2024-09-26 Jonas Nasimzada , Jens Kleesiek , Ken Herrmann , Alina Roitberg , Constantin Seibold

Deepfake is a technology dedicated to creating highly realistic facial images and videos under specific conditions, which has significant application potential in fields such as entertainment, movie production, digital human creation, to…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Gan Pei , Jiangning Zhang , Menghan Hu , Zhenyu Zhang , Chengjie Wang , Yunsheng Wu , Guangtao Zhai , Jian Yang , Dacheng Tao

We present SONIQUE, a model for generating background music tailored to video content. Unlike traditional video-to-music generation approaches, which rely heavily on paired audio-visual datasets, SONIQUE leverages unpaired data, combining…

声音 · 计算机科学 2025-02-27 Liqian Zhang , Magdalena Fuentes

The success of deep learning based face recognition systems has given rise to serious privacy concerns due to their ability to enable unauthorized tracking of users in the digital world. Existing methods for enhancing privacy fail to…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Fahad Shamshad , Muzammal Naseer , Karthik Nandakumar

Identity-preserving video generation offers powerful tools for creative expression, allowing users to customize videos featuring their beloved characters. However, prevailing methods are typically designed and optimized for a single…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Jiahao Wang , Hualian Sheng , Sijia Cai , Yuxiao Yang , Weizhan Zhang , Caixia Yan , Bing Deng , Jieping Ye

Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual…

声音 · 计算机科学 2024-10-18 Ruiqi Li , Siqi Zheng , Xize Cheng , Ziang Zhang , Shengpeng Ji , Zhou Zhao

Generating long, cohesive video stories with consistent characters is a significant challenge for current text-to-video AI. We introduce a method that approaches video generation in a filmmaker-like manner. Instead of creating a video in…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Chayan Jain , Rishant Sharma , Archit Garg , Ishan Bhanuka , Pratik Narang , Dhruv Kumar

Automated visual story generation aims to produce stories with corresponding illustrations that exhibit coherence, progression, and adherence to characters' emotional development. This work proposes a story generation pipeline to co-create…

人工智能 · 计算机科学 2023-01-10 Yuetian Chen , Ruohua Li , Bowen Shi , Peiru Liu , Mei Si

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Shuchen Weng , Haojie Zheng , Zheng Chang , Si Li , Boxin Shi , Xinlong Wang

Personalized content filtering, such as recommender systems, has become a critical infrastructure to alleviate information overload. However, these systems merely filter existing content and are constrained by its limited diversity, making…

信息检索 · 计算机科学 2025-02-05 Yiyan Xu , Wenjie Wang , Yang Zhang , Biao Tang , Peng Yan , Fuli Feng , Xiangnan He

Controllable human video generation aims to produce realistic videos of humans with explicitly guided motions and appearances,serving as a foundation for digital humans, animation, and embodied AI.However, the scarcity of largescale,…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Yuanchen Fei , Yude Zou , Zejian Kang , Ming Li , Jiaying Zhou , Xiangru Huang

Despite their impressive visual fidelity, existing personalized image generators lack interactive control over spatial composition and scale poorly to multiple humans. To address these limitations, we present LayerComposer, an interactive…

In this work, we systematically study music generation conditioned solely on the video. First, we present a large-scale dataset comprising 360K video-music pairs, including various genres such as movie trailers, advertisements, and…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Zeyue Tian , Zhaoyang Liu , Ruibin Yuan , Jiahao Pan , Qifeng Liu , Xu Tan , Qifeng Chen , Wei Xue , Yike Guo