English
Related papers

Related papers: Secure & Personalized Music-to-Video Generation vi…

200 papers

Our goal is to create a realistic 3D facial avatar with hair and accessories using only a text description. While this challenge has attracted significant recent interest, existing methods either lack realism, produce unrealistic shapes, or…

Computer Vision and Pattern Recognition · Computer Science 2023-09-14 Hao Zhang , Yao Feng , Peter Kulits , Yandong Wen , Justus Thies , Michael J. Black

Many practices have been presented in music generation recently. While stylistic music generation using deep learning techniques has became the main stream, these models still struggle to generate music with high musicality, different…

Sound · Computer Science 2021-05-12 Shuqi Dai , Xichu Ma , Ye Wang , Roger B. Dannenberg

We introduce FaceCam, a system that generates video under customizable camera trajectories for monocular human portrait video input. Recent camera control approaches based on large video-generation models have shown promising progress but…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Weijie Lyu , Ming-Hsuan Yang , Zhixin Shu

High-fidelity text-to-image diffusion models have revolutionized visual content generation, but their widespread use raises significant ethical concerns, including intellectual property protection and the misuse of synthetic media. To…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Yunzhuo Chen , Naveed Akhtar , Nur Al Hasan Haldar , Ajmal Mian

We propose a new task named Audio-driven Per-formance Video Generation (APVG), which aims to synthesizethe video of a person playing a certain instrument guided bya given music audio clip. It is a challenging task to gener-ate the…

Computer Vision and Pattern Recognition · Computer Science 2020-11-06 Hao Zhu , Yi Li , Feixia Zhu , Aihua Zheng , Ran He

The field of video generation has expanded significantly in recent years, with controllable and compositional video generation garnering considerable interest. Most methods rely on leveraging annotations such as text, objects' bounding…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Aram Davtyan , Sepehr Sameni , Björn Ommer , Paolo Favaro

This study investigates identity-preserving image synthesis, an intriguing task in image generation that seeks to maintain a subject's identity while adding a personalized, stylistic touch. Traditional methods, such as Textual Inversion and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Yuxuan Yan , Chi Zhang , Rui Wang , Yichao Zhou , Gege Zhang , Pei Cheng , Gang Yu , Bin Fu

Human motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Haiwei Xue , Xiangyang Luo , Zhanghao Hu , Xin Zhang , Xunzhi Xiang , Yuqin Dai , Jianzhuang Liu , Zhensong Zhang , Minglei Li , Jian Yang , Fei Ma , Zhiyong Wu , Changpeng Yang , Zonghong Dai , Fei Richard Yu

In human-centric content generation, the pre-trained text-to-image models struggle to produce user-wanted portrait images, which retain the identity of individuals while exhibiting diverse expressions. This paper introduces our efforts…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Renshuai Liu , Bowen Ma , Wei Zhang , Zhipeng Hu , Changjie Fan , Tangjie Lv , Yu Ding , Xuan Cheng

With the advent of personalized generation models, users can more readily create images resembling existing content, heightening the risk of violating portrait rights and intellectual property (IP). Traditional post-hoc detection and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Runyi Li , Xuanyu Zhang , Zhipei Xu , Yongbing Zhang , Jian Zhang

Customized generative text-to-image models have the ability to produce images that closely resemble a given subject. However, in the context of generating advertising images for e-commerce scenarios, it is crucial that the generated…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Youze Xue , Binghui Chen , Yifeng Geng , Xuansong Xie , Jiansheng Chen , Hongbing Ma

There is strong interest in the generation of synthetic video imagery of people talking for various purposes, including entertainment, communication, training, and advertisement. With the development of deep fake generation models,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Qiaomu Miao , Sinhwa Kang , Stacy Marsella , Steve DiPaola , Chao Wang , Ari Shapiro

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

Sound · Computer Science 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

Personalized text-to-image generation aims to create images tailored to user-defined concepts and textual descriptions. Balancing the fidelity of the learned concept with its ability for generation in various contexts presents a significant…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Vera Soboleva , Maksim Nakhodnov , Aibek Alanov

Dynamic objects in our physical 4D (3D + time) world are constantly evolving, deforming, and interacting with other objects, leading to diverse 4D scene dynamics. In this paper, we present a universal generative pipeline, CHORD, for…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Yanzhe Lyu , Chen Geng , Karthik Dharmarajan , Yunzhi Zhang , Hadi Alzayer , Shangzhe Wu , Jiajun Wu

The recent advances in deep learning have made it possible to generate photo-realistic images by using neural networks and even to extrapolate video frames from an input video clip. In this paper, for the sake of both furthering this…

Computer Vision and Pattern Recognition · Computer Science 2018-08-10 Lijie Fan , Wenbing Huang , Chuang Gan , Junzhou Huang , Boqing Gong

Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchronization, often overlooking aspects such as visual quality,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Fatemeh Nazarieh , Zhenhua Feng , Diptesh Kanojia , Muhammad Awais , Josef Kittler

Text-to-audio (TTA) generation is a recent popular problem that aims to synthesize general audio given text descriptions. Previous methods utilized latent diffusion models to learn audio embedding in a latent space with text embedding as…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Shentong Mo , Jing Shi , Yapeng Tian

Music-to-music-video generation is a challenging task due to the intrinsic differences between the music and video modalities. The advent of powerful text-to-video diffusion models has opened a promising pathway for music-video (MV)…

Sound · Computer Science 2025-03-17 Zhuoyuan Mao , Mengjie Zhao , Qiyu Wu , Zhi Zhong , Wei-Hsiang Liao , Hiromi Wakaki , Yuki Mitsufuji

The utilization of deep learning techniques in generating various contents (such as image, text, etc.) has become a trend. Especially music, the topic of this paper, has attracted widespread attention of countless researchers.The whole…

Sound · Computer Science 2020-11-16 Shulei Ji , Jing Luo , Xinyu Yang