中文
相关论文

相关论文: RealityTalk: Real-Time Speech-Driven Augmented Pre…

200 篇论文

Audience reactions can considerably enhance live experiences; conversely, in anytime-anywhere augmented reality (AR) experiences, large crowds of people might not always be available to congregate. To get closer to simulating live events…

人机交互 · 计算机科学 2025-11-05 You-Jin Kim , Misha Sra , Tobias Höllerer

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

Real-time video generation via diffusion is essential for building general-purpose multimodal interactive AI systems. However, the simultaneous denoising of all video frames with bidirectional attention via an iterative process in diffusion…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Ethan Chern , Zhulin Hu , Bohao Tang , Jiadi Su , Steffi Chern , Zhijie Deng , Pengfei Liu

This paper introduces Teachable Reality, an augmented reality (AR) prototyping tool for creating interactive tangible AR applications with arbitrary everyday objects. Teachable Reality leverages vision-based interactive machine teaching…

人机交互 · 计算机科学 2023-02-23 Kyzyl Monteiro , Ritik Vatsal , Neil Chulpongsatorn , Aman Parnami , Ryo Suzuki

Realistic visual simulations are omnipresent, yet their creation requires computing time, rendering, and expert animation knowledge. Open-vocabulary visual effects generation from text inputs emerges as a promising solution that can unlock…

图形学 · 计算机科学 2026-01-01 Luca Collorone , Mert Kiray , Indro Spinelli , Fabio Galasso , Benjamin Busam

Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Peng Chen , Xiaobao Wei , Yi Yang , Naiming Yao , Hui Chen , Feng Tian

Augmented video presentation tools provide a natural way for presenters to interact with their content, resulting in engaging experiences for remote audiences, such as when a presenter uses hand gestures to manipulate and direct attention…

人机交互 · 计算机科学 2024-06-27 Temiloluwa Femi-Gege , Matthew Brehmer , Jian Zhao

Video conferencing has become central to professional collaboration, yet most platforms offer limited support for deaf, hard-of-hearing, and multilingual users. The World Health Organisation estimates that over 430 million people worldwide…

计算工程、金融与科学 · 计算机科学 2026-04-08 Nikolaos D. Tantaroudas , Andrew J. McCracken , Ilias Karachalios , Evangelos Papatheou

Head-worn augmented reality (AR) allows audiences to be immersed and engaged in stories told by live presenters. While presenters may also be in AR to have the same level of immersion and awareness as their audience, this symmetric…

人机交互 · 计算机科学 2025-03-18 Matt Gottsacker , Mengyu Chen , David Saffo , Feiyu Lu , Benjamin Lee , Blair MacIntyre

Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Xuangeng Chu , Nabarun Goswami , Ziteng Cui , Hanqin Wang , Tatsuya Harada

Current AI writing tools, which rely on text prompts, poorly support the spatial and interactive nature of storytelling where ideas emerge from direct manipulation and play. We present PlayWrite, a mixed-reality system where users author…

人机交互 · 计算机科学 2026-03-04 Esen K. Tütüncü , Qian Zhou , Frederik Brudy , George Fitzmaurice , Fraser Anderson

In this position paper, we propose researching the combination of Augmented Reality (AR) and Artificial Intelligence (AI) to support conversations, inspired by the interfaces of dialogue systems commonly found in videogames. AR-capable…

人机交互 · 计算机科学 2025-03-10 Julián Méndez , Marc Satkowski

Generating realistic talking faces is a complex and widely discussed task with numerous applications. In this paper, we present DiffTalker, a novel model designed to generate lifelike talking faces through audio and landmark co-driving.…

计算机视觉与模式识别 · 计算机科学 2023-09-15 Zipeng Qi , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

Incorporating accurate physics-based simulation into interactive design tools is challenging. However, adding the physics accurately becomes crucial to several emerging technologies. For example, in virtual/augmented reality (VR/AR) videos,…

人机交互 · 计算机科学 2017-12-05 Dingzeyu Li

Long-term, open-domain dialogue capabilities are essential for chatbots aiming to recall past interactions and demonstrate emotional intelligence (EI). Yet, most existing research relies on synthetic, LLM-generated data, leaving open…

计算与语言 · 计算机科学 2025-02-20 Dong-Ho Lee , Adyasha Maharana , Jay Pujara , Xiang Ren , Francesco Barbieri

We present UniTalk, a novel dataset specifically designed for the task of active speaker detection, emphasizing challenging scenarios to enhance model generalization. Unlike previously established benchmarks such as AVA, which predominantly…

While Cave Automatic Virtual Environment (CAVE) systems have long enabled room-scale virtual reality and various kinds of interactivity, their content has largely remained predetermined. We present \textit{Storycaster}, a generative AI CAVE…

人机交互 · 计算机科学 2026-02-10 Naisha Agarwal , Judith Amores , Andrew D. Wilson

Task-oriented conversational agents rely on semantic parsers to translate natural language to formal representations. In this paper, we propose the design and rationale of the ThingTalk formal representation, and how the design improves the…

编程语言 · 计算机科学 2022-03-25 Monica S. Lam , Giovanni Campagna , Mehrad Moradshahi , Sina J. Semnani , Silei Xu

Communication is the most useful tool to impart knowledge, understand ideas, clarify thoughts and expressions, organize plan and manage every single day-to-day activity. Although there are different modes of communication, physical barrier…

计算机视觉与模式识别 · 计算机科学 2019-09-23 Kumar Mridul , M. Ramanathan , Kunal Ahirwar , Mansi Sharma

Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchronization, often overlooking aspects such as visual quality,…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Fatemeh Nazarieh , Zhenhua Feng , Diptesh Kanojia , Muhammad Awais , Josef Kittler