English
Related papers

Related papers: RealityTalk: Real-Time Speech-Driven Augmented Pre…

200 papers

We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or points for creating…

Computer Vision and Pattern Recognition · Computer Science 2021-02-26 Yang Zhou , Xintong Han , Eli Shechtman , Jose Echevarria , Evangelos Kalogerakis , Dingzeyu Li

Traditional lecture videos offer flexibility but lack mechanisms for real-time clarification, forcing learners to search externally when confusion arises. Recent advances in large language models and neural avatars provide new opportunities…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Md Zabirul Islam , Md Motaleb Hossen Manik , Ge Wang

We present a virtual reality (VR) environment featuring conversational avatars powered by a locally-deployed LLM, integrated with automatic speech recognition (ASR), text-to-speech (TTS), and lip-syncing. Through a pilot study, we explored…

Human-Computer Interaction · Computer Science 2025-01-08 Mykola Maslych , Christian Pumarada , Amirpouya Ghasemaghaei , Joseph J. LaViola

Storyline visualizations are an effective means to present the evolution of plots and reveal the scenic interactions among characters. However, the design of storyline visualizations is a difficult task as users need to balance between…

Human-Computer Interaction · Computer Science 2020-09-02 Tan Tang , Renzhong Li , Xinke Wu , Shuhan Liu , Johannes Knittel , Steffen Koch , Thomas Ertl , Lingyun Yu , Peiran Ren , Yingcai Wu

Understanding transcripts of immersive multimodal conversations is challenging because speakers frequently rely on visual context and non-verbal cues, such as gestures and visual attention, which are not captured in speech alone. This lack…

Human-Computer Interaction · Computer Science 2025-09-11 Riccardo Bovo , Frederik Brudy , George Fitzmaurice , Fraser Anderson

This paper presents a complete hardware and software pipeline for real-time speech enhancement in noisy and reverberant conditions. The device consists of a microphone array and a camera mounted on eyeglasses, connected to an embedded…

Current conversational AI systems aim to understand a set of pre-designed requests and execute related actions, which limits them to evolve naturally and adapt based on human interactions. Motivated by how children learn their first…

Augmented reality (AR) has revolutionized the video game industry by providing interactive, three-dimensional visualization. Interestingly, AR technology has only been sparsely used in scientific visualization. This is, at least in part,…

Graphics · Computer Science 2022-08-04 Mrudang Mathur , Josef M. Brozovich , Manuel K. Rausch

Augmented Reality (AR) is a powerful perceptual technology that can alter what users see, hear, feel, and experience throughout their daily lives. When combined with the speed and flexibility of context-aware generative AI, the power is…

Computers and Society · Computer Science 2026-01-28 Louis Rosenberg

Current talking avatars mostly generate co-speech gestures based on audio and text of the utterance, without considering the non-speaking motion of the speaker. Furthermore, previous works on co-speech gesture generation have designed…

Multimedia · Computer Science 2024-01-09 Sicheng Yang , Zunnan Xu , Haiwei Xue , Yongkang Cheng , Shaoli Huang , Mingming Gong , Zhiyong Wu

Effective presentation skills are essential in education, professional communication, and public speaking, yet learners often lack access to high-quality exemplars or personalized coaching. Existing AI tools typically provide isolated…

Human-Computer Interaction · Computer Science 2025-11-25 Sirui Chen , Jinsong Zhou , Xinli Xu , Xiaoyu Yang , Litao Guo , Ying-Cong Chen

As communications are increasingly taking place virtually, the ability to present well online is becoming an indispensable skill. Online speakers are facing unique challenges in engaging with remote audiences. However, there has been a lack…

Human-Computer Interaction · Computer Science 2023-09-12 Zeyuan Huang , Qiang He , Kevin Maher , Xiaoming Deng , Yu-Kun Lai , Cuixia Ma , Sheng-feng Qin , Yong-Jin Liu , Hongan Wang

Audio-driven facial animation presents an effective solution for animating digital avatars. In this paper, we detail the technical aspects of NVIDIA Audio2Face-3D, including data acquisition, network architecture, retargeting methodology,…

Facilitating class-wide debriefings after small-group discussions is a common strategy in ethics education. Instructor interviews revealed that effective debriefings should highlight frequently discussed themes and surface underrepresented…

Human-Computer Interaction · Computer Science 2026-01-08 Panayu Keelawat , David Barron , Kaushik Narasimhan , Daniel Manesh , Xiaohang Tang , Xi Chen , Sang Won Lee , Yan Chen

Multimodal research and applications are becoming more commonplace as Virtual Reality (VR) technology integrates different sensory feedback, enabling the recreation of real spaces in an audio-visual context. Within VR experiences, numerous…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-08 Mauricio Flores-Vargas , Enda Bates , Rachel McDonnell

Talking face generation aims to synthesize realistic speaking portraits from a single image, yet existing methods often rely on explicit optical flow and local warping, which fail to model complex global motions and cause identity drift. We…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Bo Chen , Tao Liu , Qi Chen , Xie Chen , Zilong Zheng

High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present a high-fidelity, real-time conversational digital human…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Hongbin Huang , Junwei Li , Tianxin Xie , Zhuang Li , Cekai Weng , Yaodong Yang , Yue Luo , Li Liu , Jing Tang , Zhijing Shao , Zeyu Wang

In this work, we propose a framework that creates a lively virtual dynamic scene with contextual motions of multiple humans. Generating multi-human contextual motion requires holistic reasoning over dynamic relationships among human-human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Donggeun Lim , Jinseok Bae , Inwoo Hwang , Seungmin Lee , Hwanhee Lee , Young Min Kim

High-fidelity, AI-based simulated classroom systems enable teachers to rehearse effective teaching strategies. However, dialogue-oriented open-ended conversations such as teaching a student about scale factors can be difficult to model.…

Human-Computer Interaction · Computer Science 2021-12-07 Debajyoti Datta , Maria Phillips , James P Bywater , Jennifer Chiu , Ginger S. Watson , Laura E. Barnes , Donald E Brown

In Augmented Reality (AR), virtual objects interact with real objects. However, the lack of physicality of virtual objects leads to the absence of natural sonic interactions. When virtual and real objects collide, either no sound or a…

Human-Computer Interaction · Computer Science 2025-08-05 Laura Schütz , Sasan Matinfar , Ulrich Eck , Daniel Roth , Nassir Navab