English
Related papers

Related papers: Talking Slide Avatars: Open-Source Multimodal Comm…

200 papers

This paper presents VITA (Virtual Teaching Assistants), an adaptive distributed learning (ADL) platform that embeds a large language model (LLM)-powered chatbot (BotCaptain) to provide dialogic support, interoperable analytics, and…

Computers and Society · Computer Science 2025-09-26 Fadjimata I Anaroua , Qing Li , Yan Tang , Hong P. Liu

Multimedia learning using text and images has been shown to improve learning outcomes compared to text-only instruction. But conversational AI systems in education predominantly rely on text-based interactions while multimodal conversations…

Human-Computer Interaction · Computer Science 2025-04-22 Karan Taneja , Anjali Singh , Ashok K. Goel

Accountable Talk theory has been widely adopted to analyze classroom discourse and is increasingly used to annotate tutoring interactions. In particular, the TalkMoves codebook, grounded in Accountable Talk theory, is commonly used to label…

The scarcity of high-quality, multimodal training data severely hinders the creation of lifelike avatar animations for conversational AI in virtual environments. Existing datasets often lack the intricate synchronization between speech,…

Artificial Intelligence · Computer Science 2024-10-23 Saif Punjwani , Larry Heck

Fine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT. Scaling the diversity and quality of such data, although straightforward, stands a great chance of…

Computation and Language · Computer Science 2023-05-24 Ning Ding , Yulin Chen , Bokai Xu , Yujia Qin , Zhi Zheng , Shengding Hu , Zhiyuan Liu , Maosong Sun , Bowen Zhou

Recent advances in language models have achieved significant progress. GPT-4o, as a new milestone, has enabled real-time conversations with humans, demonstrating near-human natural fluency. Such human-computer interaction necessitates…

Artificial Intelligence · Computer Science 2024-11-06 Zhifei Xie , Changqiao Wu

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have achieved notable…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yihong Lin , Zhaoxin Fan , Xianjia Wu , Lingyu Xiong , Liang Peng , Xiandong Li , Wenxiong Kang , Songju Lei , Huang Xu

The lengthy monologue-style online lectures cause learners to lose engagement easily. Designing lectures in a "vicarious dialogue" format can foster learners' cognitive activities more than monologue-style. However, designing online…

Human-Computer Interaction · Computer Science 2024-04-11 Seulgi Choi , Hyewon Lee , Yoonjoo Lee , Juho Kim

Human speech is often accompanied by hand and arm gestures. Given audio speech input, we generate plausible gestures to go along with the sound. Specifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a…

Computer Vision and Pattern Recognition · Computer Science 2019-06-11 Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , Jitendra Malik

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (AR) assistance,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Lavisha Aggarwal , Vikas Bahirwani , Andrea Colaco

The generation and editing of audio-conditioned talking portraits guided by multimodal inputs, including text, images, and videos, remains under explored. In this paper, we present SkyReels-Audio, a unified framework for synthesizing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zhengcong Fei , Hao Jiang , Di Qiu , Baoxuan Gu , Youqiang Zhang , Jiahua Wang , Jialin Bai , Debang Li , Mingyuan Fan , Guibin Chen , Yahui Zhou

The globalization of education and rapid growth of online learning have made localizing educational content a critical challenge. Lecture materials are inherently multimodal, combining spoken audio with visual slides, which requires systems…

Computation and Language · Computer Science 2026-02-24 Sai Koneru , Fabian Retkowski , Christian Huber , Lukas Hilgert , Seymanur Akti , Enes Yavuz Ugan , Alexander Waibel , Jan Niehues

Preparing high-quality instructional materials remains a labor-intensive process that often requires extensive coordination among teaching faculty, instructional designers, and teaching assistants. In this work, we present Instructional…

Artificial Intelligence · Computer Science 2026-02-03 Huaiyuan Yao , Wanpeng Xu , Justin Turnau , Nadia Kellam , Hua Wei

Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current methods still face challenges in maintaining lip-sync…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Bin Liu , Zhixiang Xiong , Zhifen He , Bo Li

Recently abstractive spoken language summarization raises emerging research interest, and neural sequence-to-sequence approaches have brought significant performance improvement. However, summarizing long meeting transcripts remains…

Computation and Language · Computer Science 2021-09-01 Zhengyuan Liu , Nancy F. Chen

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

Modeling the reactive tempo of human conversation remains difficult because most audio-visual datasets portray isolated speakers delivering short monologues. We introduce \textbf{Face-to-Face with Jimmy Fallon (F2F-JF)}, a 70-hour, 14k-clip…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ernie Chu , Vishal M. Patel

The integration of artificial intelligence (AI) into video lecture production has the potential to transform higher education by streamlining content creation and enhancing accessibility. This paper investigates a semi automated workflow…

Human-Computer Interaction · Computer Science 2025-11-27 Dengsheng Zhang

We present RealityTalk, a system that augments real-time live presentations with speech-driven interactive virtual elements. Augmented presentations leverage embedded visuals and animation for engaging and expressive storytelling. However,…

Human-Computer Interaction · Computer Science 2022-08-15 Jian Liao , Adnan Karim , Shivesh Jadon , Rubaiat Habib Kazi , Ryo Suzuki

Narrative visualization is a powerful communicative tool that can take on various formats such as interactive articles, slideshows, and data videos. These formats each have their strengths and weaknesses, but existing authoring tools only…

Human-Computer Interaction · Computer Science 2022-05-23 Matthew Conlen , Jeffrey Heer