中文
相关论文

相关论文: FlowPortrait: Reinforcement Learning for Audio-Dri…

200 篇论文

Audio-driven 3D face animation is increasingly vital in live streaming and augmented reality applications. While remarkable progress has been observed, most existing approaches are designed for specific individuals with predefined speaking…

图形学 · 计算机科学 2024-08-20 Xukun Zhou , Fengxin Li , Ziqiao Peng , Kejian Wu , Jun He , Biao Qin , Zhaoxin Fan , Hongyan Liu

Audio driven talking head synthesis is a challenging task that attracts increasing attention in recent years. Although existing methods based on 2D landmarks or 3D face models can synthesize accurate lip synchronization and rhythmic head…

计算机视觉与模式识别 · 计算机科学 2022-10-10 Yichen Han , Ya Li , Yingming Gao , Jinlong Xue , Songpo Wang , Lei Yang

Text-to-motion generation has advanced with diffusion- and flow-based generative models, yet supervised pretraining remains insufficient to align models with high-level objectives such as semantic consistency, realism, and human preference.…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Xiaofeng Tan , Wanjiang Weng , Hongsong Wang , Fang Zhao , Xin Geng , Liang Wang

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

多媒体 · 计算机科学 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

2D portrait animation has experienced significant advancements in recent years. Much research has utilized the prior knowledge embedded in large generative diffusion models to enhance high-quality image manipulation. However, most methods…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Xinya Ji , Gaspard Zoss , Prashanth Chandran , Lingchen Yang , Xun Cao , Barbara Solenthaler , Derek Bradley

We propose Expotion (Facial Expression and Motion Control for Multimodal Music Generation), a generative model leveraging multimodal visual controls - specifically, human facial expressions and upper-body motion - as well as text prompts to…

声音 · 计算机科学 2025-07-08 Fathinah Izzati , Xinyue Li , Gus Xia

Recent text-to-image systems face limitations in handling multimodal inputs and complex reasoning tasks. We introduce MindOmni, a unified multimodal large language model that addresses these challenges by incorporating reasoning generation…

人工智能 · 计算机科学 2025-06-12 Yicheng Xiao , Lin Song , Yukang Chen , Yingmin Luo , Yuxin Chen , Yukang Gan , Wei Huang , Xiu Li , Xiaojuan Qi , Ying Shan

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current…

计算机视觉与模式识别 · 计算机科学 2025-01-10 Huaize Liu , Wenzhang Sun , Donglin Di , Shibo Sun , Jiahui Yang , Changqing Zou , Hujun Bao

Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization, they often struggle with temporal consistency, identity…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Fatemeh Nazarieh , Zhenhua Feng , Diptesh Kanojia , Muhammad Awais , Josef Kittler

Vision-Language Models (VLMs) have recently shown promising advancements in sequential decision-making tasks through task-specific fine-tuning. However, common fine-tuning methods, such as Supervised Fine-Tuning (SFT) and Reinforcement…

计算与语言 · 计算机科学 2025-03-26 Haoqiang Kang , Enna Sachdeva , Piyush Gupta , Sangjae Bae , Kwonjoon Lee

An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views to learn better representations. This is further…

机器学习 · 计算机科学 2019-03-04 Nils Holzenberger , Shruti Palaskar , Pranava Madhyastha , Florian Metze , Raman Arora

Cross-modality generation is an emerging topic that aims to synthesize data in one modality based on information in a different modality. In this paper, we consider a task of such: given an arbitrary audio speech and one lip image of…

计算机视觉与模式识别 · 计算机科学 2018-05-23 Lele Chen , Zhiheng Li , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

Audio-driven human animation models often suffer from identity drift during temporal autoregressive generation, where characters gradually lose their identity over time. One solution is to generate keyframes as intermediate temporal anchors…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Junyoung Seo , Rodrigo Mira , Alexandros Haliassos , Stella Bounareli , Honglie Chen , Linh Tran , Seungryong Kim , Zoe Landgraf , Jie Shen

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

Speech-driven facial video generation has been a complex problem due to its multi-modal aspects namely audio and video domain. The audio comprises lots of underlying features such as expression, pitch, loudness, prosody(speaking style) and…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall

Synthesizing realistic videos according to a given speech is still an open challenge. Previous works have been plagued by issues such as inaccurate lip shape generation and poor image quality. The key reason is that only motions and…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Xiuzhe Wu , Pengfei Hu , Yang Wu , Xiaoyang Lyu , Yan-Pei Cao , Ying Shan , Wenming Yang , Zhongqian Sun , Xiaojuan Qi

Multimodal large language models (MLLMs) have shown remarkable performance in vision-language tasks. However, existing MLLMs are primarily trained on generic datasets, limiting their ability to reason on domain-specific visual cues such as…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Hatef Otroshi Shahreza , Sébastien Marcel

Talking head synthesis, an advanced method for generating portrait videos from a still image driven by specific content, has garnered widespread attention in virtual reality, augmented reality and game production. Recently, significant…

计算机视觉与模式识别 · 计算机科学 2024-06-19 Ming Meng , Yufei Zhao , Bo Zhang , Yonggui Zhu , Weimin Shi , Maxwell Wen , Zhaoxin Fan

We present a novel approach for synthesizing 3D facial motions from audio sequences using key motion embeddings. Despite recent advancements in data-driven techniques, accurately mapping between audio signals and 3D facial meshes remains…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Zhihao Xu , Shengjie Gong , Jiapeng Tang , Lingyu Liang , Yining Huang , Haojie Li , Shuangping Huang

Recent advances in neural portrait animation have demonstrated remarked potential for applications in virtual avatars, telepresence, and digital content creation. However, traditional explicit warping approaches often struggle with accurate…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Shihang Li , Zhiqiang Gong , Minming Ye , Yue Gao , Wen Yao