English
Related papers

Related papers: Talking Slide Avatars: Open-Source Multimodal Comm…

200 papers

Talking head synthesis, an advanced method for generating portrait videos from a still image driven by specific content, has garnered widespread attention in virtual reality, augmented reality and game production. Recently, significant…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Ming Meng , Yufei Zhao , Bo Zhang , Yonggui Zhu , Weimin Shi , Maxwell Wen , Zhaoxin Fan

Audio-driven 3D face animation is increasingly vital in live streaming and augmented reality applications. While remarkable progress has been observed, most existing approaches are designed for specific individuals with predefined speaking…

Graphics · Computer Science 2024-08-20 Xukun Zhou , Fengxin Li , Ziqiao Peng , Kejian Wu , Jun He , Biao Qin , Zhaoxin Fan , Hongyan Liu

Recent work in open-domain conversational agents has demonstrated that significant improvements in model engagingness and humanness metrics can be achieved via massive scaling in both pre-training data and model size (Adiwardana et al.,…

Computation and Language · Computer Science 2020-10-05 Kurt Shuster , Eric Michael Smith , Da Ju , Jason Weston

Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Chaoyou Fu , Haojia Lin , Xiong Wang , Yi-Fan Zhang , Yunhang Shen , Xiaoyu Liu , Haoyu Cao , Zuwei Long , Heting Gao , Ke Li , Long Ma , Xiawu Zheng , Rongrong Ji , Xing Sun , Caifeng Shan , Ran He

Talking-head avatars are increasingly adopted in educational technology to deliver content with social presence and improved engagement. However, many recent talking-head generation (THG) methods rely on GPU-centric neural rendering, large…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Vineet Kumar Rakesh , Ahana Bhattacharjee , Soumya Mazumdar , Tapas Samanta , Hemendra Kumar Pandey , Amitabha Das , Sarbajit Pal

Providing timely and actionable feedback on oral presentation slides is challenging in higher education, particularly in large classes where teachers cannot realistically deliver detailed formative feedback before students present. This…

Human-Computer Interaction · Computer Science 2026-05-07 Alvaro Becerra , Diego Gomez , Ruth Cobos

Speech-driven 3D facial animation is challenging due to the scarcity of large-scale visual-audio datasets despite extensive research. Most prior works, typically focused on learning regression models on a small dataset using the method of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Inkyu Park , Jaewoong Cho

We propose VLOGGER, a method for audio-driven human video generation from a single input image of a person, which builds on the success of recent generative diffusion models. Our method consists of 1) a stochastic human-to-3d-motion…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Enric Corona , Andrei Zanfir , Eduard Gabriel Bazavan , Nikos Kolotouros , Thiemo Alldieck , Cristian Sminchisescu

With children talking to smart-speakers, smart-phones and even smart-microwaves daily, it is increasingly important to educate students on how these agents work-from underlying mechanisms to societal implications. Researchers are developing…

Computers and Society · Computer Science 2020-09-15 Jessica Van Brummelen , Tommy Heng , Viktoriya Tabunshchyk

Virtual reality (VR) and interactive 3D visualization systems have enhanced educational experiences and environments, particularly in complicated subjects such as anatomy education. VR-based systems surpass the potential limitations of…

Current talking face generation methods mainly focus on speech-lip synchronization. However, insufficient investigation on the facial talking style leads to a lifeless and monotonous avatar. Most previous works fail to imitate expressive…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Liyang Chen , Zhiyong Wu , Runnan Li , Weihong Bao , Jun Ling , Xu Tan , Sheng Zhao

We introduce DrawTalking, an approach to building and controlling interactive worlds by sketching and speaking while telling stories. It emphasizes user control and flexibility, and gives programming-like capability without requiring code.…

Human-Computer Interaction · Computer Science 2024-08-07 Karl Toby Rosenberg , Rubaiat Habib Kazi , Li-Yi Wei , Haijun Xia , Ken Perlin

We propose MAViD, a novel Multimodal framework for Audio-Visual Dialogue understanding and generation. Existing approaches primarily focus on non-interactive systems and are limited to producing constrained and unnatural human speech. The…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Youxin Pang , Jiajun Liu , Lingfeng Tan , Yong Zhang , Feng Gao , Xiang Deng , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

Audio-Visual Scene-Aware Dialog (AVSD) is a task to generate responses when chatting about a given video, which is organized as a track of the 8th Dialog System Technology Challenge (DSTC8). To solve the task, we propose a universal…

Computation and Language · Computer Science 2020-02-04 Zekang Li , Zongjia Li , Jinchao Zhang , Yang Feng , Cheng Niu , Jie Zhou

Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistants. However, most existing approaches are limited to offline processing of complete audio…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Bohong Chen , Yumeng Li , Yinglin Xu , Youyi Zheng , Yanlin Weng , Kun Zhou

Large Language Models have demonstrated remarkable capabilities in open-domain dialogues. However, current methods exhibit suboptimal performance in service dialogues, as they rely on noisy, low-quality human conversation data. This…

Computation and Language · Computer Science 2026-05-06 Yuqin Dai , Ning Gao , Wei Zhang , Jie Wang , Zichen Luo , Jinpeng Wang , Yujie Wang , Ruiyuan Wu , Chaozheng Wang

This brief literature review studies the problem of audiovisual speech synthesis, which is the problem of generating an animated talking head given a text as input. Due to the high complexity of this problem, we approach it as the…

Sound · Computer Science 2021-03-09 Efthymios Georgiou , Athanasios Katsamanis

Audio-driven human video generation has achieved remarkable success in monologue scenarios, largely driven by advancements in powerful video generation foundation models. Moving beyond monologues, authentic human communication is inherently…

Artificial Intelligence · Computer Science 2026-04-14 Yuzhe Weng , Haotian Wang , Xinyi Yu , Xiaoyan Wu , Haoran Xu , Shan He , Jun Du

Speech-driven 3D face animation technique, extending its applications to various multimedia fields. Previous research has generated promising realistic lip movements and facial expressions from audio signals. However, traditional regression…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Ziqiao Peng , Yihao Luo , Yue Shi , Hao Xu , Xiangyu Zhu , Jun He , Hongyan Liu , Zhaoxin Fan

Academic presentation videos have become an essential medium for research communication, yet producing them remains highly labor-intensive, often requiring hours of slide design, recording, and editing for a short 2 to 10 minutes video.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Zeyu Zhu , Kevin Qinghong Lin , Mike Zheng Shou