中文
相关论文

相关论文: Speech2Video: Cross-Modal Distillation for Speech …

200 篇论文

Customized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches for personalizing text-to-video generation suffer from tackling multiple subjects, which is a more…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Zhao Wang , Aoxue Li , Lingting Zhu , Yong Guo , Qi Dou , Zhenguo Li

Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic…

Understanding and analyzing video actions are essential for producing insightful and contextualized descriptions, especially for video-based applications like intelligent monitoring and autonomous systems. The proposed work introduces a…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Lakshita Agarwal , Bindu Verma

Audio-driven talking head animation is a challenging research topic with many real-world applications. Recent works have focused on creating photo-realistic 2D animation, while learning different talking or singing styles remains an open…

计算机视觉与模式识别 · 计算机科学 2023-03-23 Trong-Thang Pham , Nhat Le , Tuong Do , Hung Nguyen , Erman Tjiputra , Quang D. Tran , Anh Nguyen

The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Chenxu Zhang , Chao Wang , Jianfeng Zhang , Hongyi Xu , Guoxian Song , You Xie , Linjie Luo , Yapeng Tian , Xiaohu Guo , Jiashi Feng

The objective of this paper is to learn representations of speaker identity without access to manually annotated data. To do so, we develop a self-supervised learning objective that exploits the natural cross-modal synchrony between faces…

音频与语音处理 · 电气工程与系统科学 2020-05-05 Arsha Nagrani , Joon Son Chung , Samuel Albanie , Andrew Zisserman

In recent years, with the realistic generation results and a wide range of personalized applications, diffusion-based generative models gain huge attention in both visual and audio generation areas. Compared to the considerable advancements…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Shiqi Yang , Zhi Zhong , Mengjie Zhao , Shusuke Takahashi , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

Talking head synthesis, an advanced method for generating portrait videos from a still image driven by specific content, has garnered widespread attention in virtual reality, augmented reality and game production. Recently, significant…

计算机视觉与模式识别 · 计算机科学 2024-06-19 Ming Meng , Yufei Zhao , Bo Zhang , Yonggui Zhu , Weimin Shi , Maxwell Wen , Zhaoxin Fan

As artificial intelligence-generated content (AIGC) continues to evolve, video-to-audio (V2A) generation has emerged as a key area with promising applications in multimedia editing, augmented reality, and automated content creation. While…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Yuhuan You , Xihong Wu , Tianshu Qu

This work proposes a novel method to generate realistic talking head videos using audio and visual streams. We animate a source image by transferring head motion from a driving video using a dense motion field generated using learnable…

计算机视觉与模式识别 · 计算机科学 2022-10-07 Madhav Agarwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Zhongjian Wang , Peng Zhang , Jinwei Qi , Guangyuan Wang , Chaonan Ji , Sheng Xu , Bang Zhang , Liefeng Bo

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

图形学 · 计算机科学 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang

Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Yuchi Wang , Junliang Guo , Jianhong Bai , Runyi Yu , Tianyu He , Xu Tan , Xu Sun , Jiang Bian

Anticipating future actions in a video is useful for many autonomous and assistive technologies. Most prior action anticipation work treat this as a vision modality problem, where the models learn the task information primarily from the…

计算机视觉与模式识别 · 计算机科学 2023-02-22 Sayontan Ghosh , Tanvi Aggarwal , Minh Hoai , Niranjan Balasubramanian

This paper proposes visual-text to speech (vTTS), a method for synthesizing speech from visual text (i.e., text as an image). Conventional TTS converts phonemes or characters into discrete symbols and synthesizes a speech waveform from…

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Jiadong Liang , Feng Lu

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions…

机器学习 · 计算机科学 2021-01-15 Simon Alexanderson , Éva Székely , Gustav Eje Henter , Taras Kucherenko , Jonas Beskow

Video generation has achieved remarkable progress with the introduction of diffusion models, which have significantly improved the quality of generated videos. However, recent research has primarily focused on scaling up model training,…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Chenyang Si , Weichen Fan , Zhengyao Lv , Ziqi Huang , Yu Qiao , Ziwei Liu

This work addresses the problem of generating 3D holistic body motions from human speech. Given a speech recording, we synthesize sequences of 3D body poses, hand gestures, and facial expressions that are realistic and diverse. To achieve…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Hongwei Yi , Hualin Liang , Yifei Liu , Qiong Cao , Yandong Wen , Timo Bolkart , Dacheng Tao , Michael J. Black

Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Xuli Shen , Hua Cai , Dingding Yu , Weilin Shen , Qing Xu , Xiangyang Xue