English
Related papers

Related papers: RealityTalk: Real-Time Speech-Driven Augmented Pre…

200 papers

Audience reactions can considerably enhance live experiences; conversely, in anytime-anywhere augmented reality (AR) experiences, large crowds of people might not always be available to congregate. To get closer to simulating live events…

Human-Computer Interaction · Computer Science 2025-11-05 You-Jin Kim , Misha Sra , Tobias Höllerer

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

Real-time video generation via diffusion is essential for building general-purpose multimodal interactive AI systems. However, the simultaneous denoising of all video frames with bidirectional attention via an iterative process in diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Ethan Chern , Zhulin Hu , Bohao Tang , Jiadi Su , Steffi Chern , Zhijie Deng , Pengfei Liu

This paper introduces Teachable Reality, an augmented reality (AR) prototyping tool for creating interactive tangible AR applications with arbitrary everyday objects. Teachable Reality leverages vision-based interactive machine teaching…

Human-Computer Interaction · Computer Science 2023-02-23 Kyzyl Monteiro , Ritik Vatsal , Neil Chulpongsatorn , Aman Parnami , Ryo Suzuki

Realistic visual simulations are omnipresent, yet their creation requires computing time, rendering, and expert animation knowledge. Open-vocabulary visual effects generation from text inputs emerges as a promising solution that can unlock…

Graphics · Computer Science 2026-01-01 Luca Collorone , Mert Kiray , Indro Spinelli , Fabio Galasso , Benjamin Busam

Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Peng Chen , Xiaobao Wei , Yi Yang , Naiming Yao , Hui Chen , Feng Tian

Augmented video presentation tools provide a natural way for presenters to interact with their content, resulting in engaging experiences for remote audiences, such as when a presenter uses hand gestures to manipulate and direct attention…

Human-Computer Interaction · Computer Science 2024-06-27 Temiloluwa Femi-Gege , Matthew Brehmer , Jian Zhao

Video conferencing has become central to professional collaboration, yet most platforms offer limited support for deaf, hard-of-hearing, and multilingual users. The World Health Organisation estimates that over 430 million people worldwide…

Computational Engineering, Finance, and Science · Computer Science 2026-04-08 Nikolaos D. Tantaroudas , Andrew J. McCracken , Ilias Karachalios , Evangelos Papatheou

Head-worn augmented reality (AR) allows audiences to be immersed and engaged in stories told by live presenters. While presenters may also be in AR to have the same level of immersion and awareness as their audience, this symmetric…

Human-Computer Interaction · Computer Science 2025-03-18 Matt Gottsacker , Mengyu Chen , David Saffo , Feiyu Lu , Benjamin Lee , Blair MacIntyre

Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Xuangeng Chu , Nabarun Goswami , Ziteng Cui , Hanqin Wang , Tatsuya Harada

Current AI writing tools, which rely on text prompts, poorly support the spatial and interactive nature of storytelling where ideas emerge from direct manipulation and play. We present PlayWrite, a mixed-reality system where users author…

Human-Computer Interaction · Computer Science 2026-03-04 Esen K. Tütüncü , Qian Zhou , Frederik Brudy , George Fitzmaurice , Fraser Anderson

In this position paper, we propose researching the combination of Augmented Reality (AR) and Artificial Intelligence (AI) to support conversations, inspired by the interfaces of dialogue systems commonly found in videogames. AR-capable…

Human-Computer Interaction · Computer Science 2025-03-10 Julián Méndez , Marc Satkowski

Generating realistic talking faces is a complex and widely discussed task with numerous applications. In this paper, we present DiffTalker, a novel model designed to generate lifelike talking faces through audio and landmark co-driving.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Zipeng Qi , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

Incorporating accurate physics-based simulation into interactive design tools is challenging. However, adding the physics accurately becomes crucial to several emerging technologies. For example, in virtual/augmented reality (VR/AR) videos,…

Human-Computer Interaction · Computer Science 2017-12-05 Dingzeyu Li

Long-term, open-domain dialogue capabilities are essential for chatbots aiming to recall past interactions and demonstrate emotional intelligence (EI). Yet, most existing research relies on synthetic, LLM-generated data, leaving open…

Computation and Language · Computer Science 2025-02-20 Dong-Ho Lee , Adyasha Maharana , Jay Pujara , Xiang Ren , Francesco Barbieri

We present UniTalk, a novel dataset specifically designed for the task of active speaker detection, emphasizing challenging scenarios to enhance model generalization. Unlike previously established benchmarks such as AVA, which predominantly…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Le Thien Phuc Nguyen , Zhuoran Yu , Khoa Quang Nhat Cao , Yuwei Guo , Tu Ho Manh Pham , Tuan Tai Nguyen , Toan Ngo Duc Vo , Lucas Poon , Soochahn Lee , Yong Jae Lee

While Cave Automatic Virtual Environment (CAVE) systems have long enabled room-scale virtual reality and various kinds of interactivity, their content has largely remained predetermined. We present \textit{Storycaster}, a generative AI CAVE…

Human-Computer Interaction · Computer Science 2026-02-10 Naisha Agarwal , Judith Amores , Andrew D. Wilson

Task-oriented conversational agents rely on semantic parsers to translate natural language to formal representations. In this paper, we propose the design and rationale of the ThingTalk formal representation, and how the design improves the…

Programming Languages · Computer Science 2022-03-25 Monica S. Lam , Giovanni Campagna , Mehrad Moradshahi , Sina J. Semnani , Silei Xu

Communication is the most useful tool to impart knowledge, understand ideas, clarify thoughts and expressions, organize plan and manage every single day-to-day activity. Although there are different modes of communication, physical barrier…

Computer Vision and Pattern Recognition · Computer Science 2019-09-23 Kumar Mridul , M. Ramanathan , Kunal Ahirwar , Mansi Sharma

Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchronization, often overlooking aspects such as visual quality,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Fatemeh Nazarieh , Zhenhua Feng , Diptesh Kanojia , Muhammad Awais , Josef Kittler