English
Related papers

Related papers: Talking Slide Avatars: Open-Source Multimodal Comm…

200 papers

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

Large Language Models (LLMs) are increasingly employed in multi-turn conversational tasks, yet their pre-training data predominantly consists of continuous prose, creating a potential mismatch between required capabilities and training…

Computation and Language · Computer Science 2025-07-09 Jing Yang Lee , Hamed Bonab , Nasser Zalmout , Ming Zeng , Sanket Lokegaonkar , Colin Lockard , Binxuan Huang , Ritesh Sarkhel , Haodong Wang

Speech-driven facial animation methods usually contain two main classes, 3D and 2D talking face, both of which attract considerable research attention in recent years. However, to the best of our knowledge, the research on 3D talking face…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Yixiang Zhuang , Baoping Cheng , Yao Cheng , Yuntao Jin , Renshuai Liu , Chengyang Li , Xuan Cheng , Jing Liao , Juncong Lin

The paper introduces AniTalker, an innovative framework designed to generate lifelike talking faces from a single portrait. Unlike existing models that primarily focus on verbal cues such as lip synchronization and fail to capture the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Tao Liu , Feilong Chen , Shuai Fan , Chenpeng Du , Qi Chen , Xie Chen , Kai Yu

In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ziqiao Peng , Yanbo Fan , Haoyu Wu , Xuan Wang , Hongyan Liu , Jun He , Zhaoxin Fan

Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchronization, often overlooking aspects such as visual quality,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Fatemeh Nazarieh , Zhenhua Feng , Diptesh Kanojia , Muhammad Awais , Josef Kittler

We introduce TalkVerse, a large-scale, open corpus for single-person, audio-driven talking video generation designed to enable fair, reproducible comparison across methods. While current state-of-the-art systems rely on closed data or…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Zhenzhi Wang , Jian Wang , Ke Ma , Dahua Lin , Bing Zhou

One-shot talking head video generation uses a source image and driving video to create a synthetic video where the source person's facial movements imitate those of the driving video. However, differences in scale between the source and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Fa-Ting Hong , Dan Xu

People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth wide open, while the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Haozhe Wu , Jia Jia , Haoyu Wang , Yishun Dou , Chao Duan , Qingshan Deng

Large Language Models (LLMs) have advanced rapidly in recent years. One application of LLMs is to support student learning in educational settings. However, prior work has shown that LLMs still struggle to answer questions accurately within…

Computation and Language · Computer Science 2026-03-19 Tu Anh Dinh , Philipp Nicolas Schumacher , Jan Niehues

When humans converse, what a speaker will say next significantly depends on what he sees. Unfortunately, existing dialogue models generate dialogue utterances only based on preceding textual contexts, and visual contexts are rarely…

Computation and Language · Computer Science 2021-06-01 Yuxian Meng , Shuhe Wang , Qinghong Han , Xiaofei Sun , Fei Wu , Rui Yan , Jiwei Li

Audio-visual speech recognition (AVSR) is a multimodal extension of automatic speech recognition (ASR), using video as a complement to audio. In AVSR, considerable efforts have been directed at datasets for facial features such as…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Hao Wang , Shuhei Kurita , Shuichiro Shimizu , Daisuke Kawahara

We present an architecture for integrating real-time, multimodal input into a computational agent's contextual model. Using a human-avatar interaction in a virtual world, we treat aligned gesture and speech as an ensemble where content may…

Human-Computer Interaction · Computer Science 2019-09-19 Nikhil Krishnaswamy , James Pustejovsky

Audio-driven 3D facial animation has several virtual humans applications for content creation and editing. While several existing methods provide solutions for speech-driven animation, precise control over content (what) and style (how) of…

Sound · Computer Science 2024-08-15 Qingju Liu , Hyeongwoo Kim , Gaurav Bharaj

Audio-driven talking head generation aims to create vivid and realistic videos from a static portrait and speech. Existing AR-based methods rely on intermediate facial representations, which limit their expressiveness and realism.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yuzhe Weng , Haotian Wang , Yuanhong Yu , Jun Du , Shan He , Xiaoyan Wu , Haoran Xu

The rapid development of large-scale models has catalyzed significant breakthroughs in the digital human domain. These advanced methodologies offer high-fidelity solutions for avatar driving and rendering, leading academia to focus on the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Youliang Zhang , Zhaoyang Li , Duomin Wang , Jiahe Zhang , Deyu Zhou , Zixin Yin , Xili Dai , Gang Yu , Xiu Li

Portrait animation aims to synthesize talking videos from a static reference face, conditioned on audio and style frame cues (e.g., emotion and head poses), while ensuring precise lip synchronization and faithful reproduction of speaking…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 He Feng , Yongjia Ma , Donglin Di , Lei Fan , Tonghua Su , Xiangqian Wu

Effective feedback is essential for refining instructional practices in mathematics education, and researchers often turn to advanced natural language processing (NLP) models to analyze classroom dialogues from multiple perspectives.…

Computation and Language · Computer Science 2025-08-05 Jannatun Naim , Jie Cao , Fareen Tasneem , Jennifer Jacobs , Brent Milne , James Martin , Tamara Sumner

Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Peng Chen , Xiaobao Wei , Yi Yang , Naiming Yao , Hui Chen , Feng Tian

The exponential growth of AI education has brought millions of learners to online platforms, yet this massive scale has simultaneously exposed critical pedagogical shortcomings. Traditional video-based instruction, while cost-effective and…

Human-Computer Interaction · Computer Science 2026-04-20 Mohammed Abraar , Raj Abhijit Dandekar , Rajat Dandekar , Sreedath Panat