English
Related papers

Related papers: Talking Slide Avatars: Open-Source Multimodal Comm…

200 papers

Online learning and academic conferences have become pervasive and essential for education and professional development, especially since the onset of pandemics. Academic presentations usually require well-designed slides that are easily…

Human-Computer Interaction · Computer Science 2023-04-26 Jiahao Weng , Xusheng Du , Haoran Xie

Massive Open Online Courses (MOOCs) platforms are becoming increasingly popular in recent years. Online learners need to watch the whole course video on MOOC platforms to learn the underlying new knowledge, which is often tedious and…

Human-Computer Interaction · Computer Science 2024-02-23 Zhiguang Zhou , Li Ye , Lihong Cai , Lei Wang , Yigang Wang , Yongheng Wang , Wei Chen , Yong Wang

Recent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion from just a single RGB image and a set of driving poses.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Alireza Javanmardi , Pragati Jaiswal , Tewodros Amberbir Habtegebrial , Christen Millerdurai , Shaoxiang Wang , Alain Pagani , Didier Stricker

State-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems primarily rely on acoustic information while disregarding additional multi-modal context. However, visual information are essential in disambiguation and adaptation. While…

Artificial Intelligence · Computer Science 2025-10-17 Supriti Sinhamahapatra , Jan Niehues

Recently, 2D speaking avatars have increasingly participated in everyday scenarios due to the fast development of facial animation techniques. However, most existing works neglect the explicit control of human bodies. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Jiazhi Guan , Quanwei Yang , Kaisiyuan Wang , Hang Zhou , Shengyi He , Zhiliang Xu , Haocheng Feng , Errui Ding , Jingdong Wang , Hongtao Xie , Youjian Zhao , Ziwei Liu

The dialogue experience with conversational agents can be greatly enhanced with multimodal and immersive interactions in virtual reality. In this work, we present an open-source architecture with the goal of simplifying the development of…

Artificial Intelligence · Computer Science 2023-08-08 Michele Yin , Gabriel Roccabruna , Abhinav Azad , Giuseppe Riccardi

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

Graphics · Computer Science 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang

Video-based programming tutorials are a popular form of tutorial used by authors to guide learners to code. Still, the interactivity of these videos is limited primarily to control video flow. There are existing works with increased…

Software Engineering · Computer Science 2022-04-20 Eng Lieh Ouh , Benjamin Kok Siew Gan , David Lo

Large language models (LLMs) have advanced virtual educators and learners, bridging NLP with AI4Education. Existing work often lacks scalability and fails to leverage diverse, large-scale course content, with limited frameworks for…

Artificial Intelligence · Computer Science 2025-09-08 Jiahuan Pei , Fanghua Ye , Xin Sun , Wentao Deng , Koen Hindriks , Junxiao Wang

Many educational organizations are employing instructional video in their pedagogy, but there is limited understanding of the possible presentation styles. In practice, the presentation style of video lectures ranges from a direct recording…

Computers and Society · Computer Science 2020-01-01 Konstantinos Chorianopoulos

In recent years, synthetic media from deepfake videos have emerged as a new interesting technology, whether that refers to cloned voices, multilingual translation models, or more recent applications of avatar tutors into higher education.…

Computers and Society · Computer Science 2026-02-24 Alexandros Gazis , Erietta Chamalidou , Nikolaos Ntaoulas , Theodoros Vavouras

Controllability, generalizability and efficiency are the major objectives of constructing face avatars represented by neural implicit field. However, existing methods have not managed to accommodate the three requirements simultaneously.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Zhiyuan Ma , Xiangyu Zhu , Guojun Qi , Zhen Lei , Lei Zhang

Audio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, expressive facial expressions, natural head pose generation, and high video quality. However, no model has yet led…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Xusen Sun , Longhao Zhang , Hao Zhu , Peng Zhang , Bang Zhang , Xinya Ji , Kangneng Zhou , Daiheng Gao , Liefeng Bo , Xun Cao

We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or points for creating…

Computer Vision and Pattern Recognition · Computer Science 2021-02-26 Yang Zhou , Xintong Han , Eli Shechtman , Jose Echevarria , Evangelos Kalogerakis , Dingzeyu Li

Different people speak with diverse personalized speaking styles. Although existing one-shot talking head methods have made significant progress in lip sync, natural facial expressions, and stable head motions, they still cannot generate…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Yifeng Ma , Suzhen Wang , Zhipeng Hu , Changjie Fan , Tangjie Lv , Yu Ding , Zhidong Deng , Xin Yu

Large-language-model assistants are suitable for explaining popular APIs, yet they falter on niche or proprietary libraries because the multi-turn dialogue data needed for fine-tuning are scarce. We present APIDA-Chat, an open-source…

Software Engineering · Computer Science 2025-10-07 Zachary Eberhart , Collin McMillan

We propose a neural talking-head video synthesis model and demonstrate its application to video conferencing. Our model learns to synthesize a talking-head video using a source image containing the target person's appearance and a driving…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Ting-Chun Wang , Arun Mallya , Ming-Yu Liu

Presentations are a primary medium for scholarly communication, yet most AI slide generators optimize the artifact (a visually plausible deck) while under-optimizing the delivery process (pacing, narrative, and presentation preparation). We…

Artificial Intelligence · Computer Science 2026-05-18 Ming Yang , Zhiwei Zhang , Jiahang Li , Haoseng Liu , Yuzheng Cai , Weiguo Zheng

This work expands further our earlier poster presentation and integration of the OpenGL Slides Framework (OGLSF) - to make presentations with real-time animated graphics where each slide is a scene with tidgets - and physical based…

Graphics · Computer Science 2010-07-06 Miao Song , Serguei A. Mokhov , Peter Grogono

Educational videos are widely used across various instructional models in higher education to support flexible and self-paced learning. However, student engagement with these videos varies significantly depending on how they are designed.…

Applications · Statistics 2025-12-19 Mohamed Tolba , Olivia Kendall , Daniel Tudball Smith , Alexander Gregg , Tony Vo , Scott Wordley