中文
相关论文

相关论文: Wan-S2V: Audio-Driven Cinematic Video Generation

200 篇论文

The importance of modeling speech articulation for high-quality audiovisual (AV) speech synthesis is widely acknowledged. Nevertheless, while state-of-the-art, data-driven approaches to facial animation can make use of sophisticated motion…

人机交互 · 计算机科学 2012-09-25 Ingmar Steiner , Korin Richmond , Slim Ouni

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input…

计算机视觉与模式识别 · 计算机科学 2023-01-03 Balamurugan Thambiraja , Ikhsanul Habibie , Sadegh Aliakbarian , Darren Cosker , Christian Theobalt , Justus Thies

Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Haojie Zheng , Yixin Yang , Siqi Yang , Shuchen Weng , Boxin Shi

Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital…

声音 · 计算机科学 2025-10-15 Tianbao Zhang , Jian Zhao , Yuer Li , Zheng Zhu , Ping Hu , Zhaoxin Fan , Wenjun Wu , Xuelong Li

The creation of a parameterized stylized character involves careful selection of numerous parameters, also known as the "avatar vectors" that can be interpreted by the avatar engine. Existing unsupervised avatar vector estimation methods…

计算机视觉与模式识别 · 计算机科学 2023-02-20 Shizun Wang , Weihong Zeng , Xu Wang , Hao Yang , Li Chen , Yi Yuan , Yunzhao Zeng , Min Zheng , Chuang Zhang , Ming Wu

High-fidelity and efficient audio-driven talking head generation has been a key research topic in computer graphics and computer vision. In this work, we study vector image based audio-driven talking head generation. Compared with directly…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Hao Hu , Xuan Wang , Jingxiang Sun , Yanbo Fan , Yu Guo , Caigui Jiang

We introduce the first method for audio-driven universal photorealistic avatar synthesis, combining a person-agnostic speech model with our novel Universal Head Avatar Prior (UHAP). UHAP is trained on cross-identity multi-view videos. In…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Kartik Teotia , Helge Rhodin , Mohit Mendiratta , Hyeongwoo Kim , Marc Habermann , Christian Theobalt

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio…

We propose an end to end deep learning approach for generating real-time facial animation from just audio. Specifically, our deep architecture employs deep bidirectional long short-term memory network and attention mechanism to discover the…

机器学习 · 计算机科学 2019-05-28 Guanzhong Tian , Yi Yuan , Yong liu

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity,…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Ziwei Zhou , Zeyuan Lai , Rui Wang , Yifan Yang , Zhen Xing , Yuqing Yang , Qi Dai , Lili Qiu , Chong Luo

Large vision and language models show strong performance in tasks like image captioning, visual question answering, and retrieval. However, challenges remain in integrating speech, text, and vision into a unified model, especially for…

多媒体 · 计算机科学 2025-07-08 Ngoc Dung Huynh , Mohamed Reda Bouadjenek , Imran Razzak , Hakim Hacid , Sunil Aryal

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-matching generative…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Peiyin Chen , Zhuowei Yang , Hui Feng , Sheng Jiang , Rui Yan

In this paper, we present JoVA, a unified framework for joint video-audio generation. Despite recent encouraging advances, existing methods face two critical limitations. First, most existing approaches can only generate ambient sounds and…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Xiaohu Huang , Hao Zhou , Qiangpeng Yang , Shilei Wen , Kai Han

Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted, single-view perspective lacks sufficient temporal and…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Zujin Guo , Zhenhui Ye , Yi Ren , Yuanming Li , Ce Chen , Zhibin Hong , Chen Change Loy

While Text-To-Video (T2V) models have advanced rapidly, they continue to struggle with generating legible and coherent text within videos. In particular, existing models often fail to render correctly even short phrases or words and…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Ziyang Liu , Kevin Valencia , Justin Cui

Whenever we speak, our voice is accompanied by facial movements and expressions. Several recent works have shown the synthesis of highly photo-realistic videos of talking faces, but they either require a source video to drive the target…

计算机视觉与模式识别 · 计算机科学 2020-11-03 Prateek Manocha , Prithwijit Guha

Audio-video generation has often relied on complex multi-stage architectures or sequential synthesis of sound and visuals. We introduce Ovi, a unified paradigm for audio-video generation that models the two modalities as a single generative…

多媒体 · 计算机科学 2025-10-03 Chetwin Low , Weimin Wang , Calder Katyal

Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes and asynchronous lips because those methods struggle to learn…

计算机视觉与模式识别 · 计算机科学 2021-12-07 Suzhen Wang , Lincheng Li , Yu Ding , Xin Yu

Music-to-Video (M2V) generation for full-length songs faces significant challenges. Existing methods produce short, disjointed clips, failing to align visuals with musical structure, beats, or lyrics, and lack temporal consistency. We…

Current audio-driven facial animation methods achieve impressive results for short videos but suffer from error accumulation and identity drift when extended to longer durations. Existing methods attempt to mitigate this through external…