English
Related papers

Related papers: Storycaster: An AI System for Immersive Room-Based…

200 papers

Generative AI (GenAI) has significantly advanced the ease and flexibility of image creation. However, it remains a challenge to precisely control spatial compositions, including object arrangement and scene conditions. To bridge this gap,…

Human-Computer Interaction · Computer Science 2025-08-12 Runlin Duan , Yuzhao Chen , Rahul Jain , Yichen Hu , Jingyu Shi , Karthik Ramani

Conversational agents (CAs) have the great potential in mitigating the clinicians' burden in screening for neurocognitive disorders among older adults. It is important, therefore, to develop CAs that can be engaging, to elicit…

Human-Computer Interaction · Computer Science 2022-02-17 Zijian Ding , Jiawen Kang , Tinky Oi Ting HO , Ka Ho Wong , Helene H. Fung , Helen Meng , Xiaojuan Ma

Text-to-audio diffusion models produce high-fidelity audio but require tens of function evaluations (NFEs), incurring multi-second latency and limited throughput. We present SoundWeaver, the first training-free, model-agnostic serving…

Sound · Computer Science 2026-03-10 Ayush Barik , Sofia Stoica , Nikhil Sarda , Arnav Kethana , Abhinav Khanduja , Muchen Xu , Fan Lai

Large Language Models show promise for AI-assisted storytelling, yet current tools often generate predictable, unoriginal narratives. To address this limitation, we present NarrativeLoom, a multi-persona co-creative system grounded in…

Human-Computer Interaction · Computer Science 2026-03-10 Yuxi Ma , Yongqian Peng , Fengyuan Yang , Siyu Zha , Chi Zhang , Zixia Jia , Zilong Zheng , Yixin Zhu

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Yiguo Jiang , Xiaodong Cun , Yong Zhang , Yudian Zheng , Fan Tang , Chi-Man Pun

How should an AI-based explanation system explain an agent's complex behavior to ordinary end users who have no background in AI? Answering this question is an active research area, for if an AI-based explanation system could effectively…

Human-Computer Interaction · Computer Science 2017-11-21 Jonathan Dodge , Sean Penney , Claudia Hilderbrand , Andrew Anderson , Margaret Burnett

We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Aggelina Chatziagapi , Louis-Philippe Morency , Hongyu Gong , Michael Zollhoefer , Dimitris Samaras , Alexander Richard

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

Computer Vision and Pattern Recognition · Computer Science 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu

There has been a recent explosion of impressive generative models that can produce high quality images (or videos) conditioned on text descriptions. However, all such approaches rely on conditional sentences that contain unambiguous…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Tanzila Rahman , Hsin-Ying Lee , Jian Ren , Sergey Tulyakov , Shweta Mahajan , Leonid Sigal

In this position paper, we propose researching the combination of Augmented Reality (AR) and Artificial Intelligence (AI) to support conversations, inspired by the interfaces of dialogue systems commonly found in videogames. AR-capable…

Human-Computer Interaction · Computer Science 2025-03-10 Julián Méndez , Marc Satkowski

Effective presentation skills are essential in education, professional communication, and public speaking, yet learners often lack access to high-quality exemplars or personalized coaching. Existing AI tools typically provide isolated…

Human-Computer Interaction · Computer Science 2025-11-25 Sirui Chen , Jinsong Zhou , Xinli Xu , Xiaoyu Yang , Litao Guo , Ying-Cong Chen

This paper presents the development of an interactive system for constructing Augmented Virtual Environments (AVEs) by fusing mobile phone images with open-source geospatial data. By integrating 2D image data with 3D models derived from…

Human-Computer Interaction · Computer Science 2025-09-19 Russell Beale , Daniel Rutter

Interactive storytelling is vital for preschooler development. While children's interactive partners have traditionally been their parents and teachers, recent advances in artificial intelligence (AI) have sparked a surge of AI-based…

Human-Computer Interaction · Computer Science 2024-09-04 Yuling Sun , Jiaju Chen , Bingsheng Yao , Jiali Liu , Dakuo Wang , Xiaojuan Ma , Yuxuan Lu , Ying Xu , Liang He

Multimodal audiovisual perception can enable new avenues for robotic manipulation, from better material classification to the imitation of demonstrations for which only audio signals are available (e.g., playing a tune by ear). However, to…

Multi-modal AI systems will likely become a ubiquitous presence in our everyday lives. A promising approach to making these systems more interactive is to embody them as agents within physical and virtual environments. At present, systems…

Artificial Intelligence (AI), especially Neural Networks (NNs), has become increasingly popular. However, people usually treat AI as a tool, focusing on improving outcome, accuracy, and performance while paying less attention to the…

Human-Computer Interaction · Computer Science 2021-11-16 Zhuoyue Lyu , Jiannan Li , Bryan Wang

For audio in augmented reality (AR), knowledge of the users' real acoustic environment is crucial for rendering virtual sounds that seamlessly blend into the environment. As acoustic measurements are usually not feasible in practical AR…

Sound · Computer Science 2024-09-24 Francesc Lluís , Nils Meyer-Kahlen

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

Storyline visualizations are an effective means to present the evolution of plots and reveal the scenic interactions among characters. However, the design of storyline visualizations is a difficult task as users need to balance between…

Human-Computer Interaction · Computer Science 2020-09-02 Tan Tang , Renzhong Li , Xinke Wu , Shuhan Liu , Johannes Knittel , Steffen Koch , Thomas Ertl , Lingyun Yu , Peiran Ren , Yingcai Wu

Immersive authoring provides an intuitive medium for users to create 3D scenes via direct manipulation in Virtual Reality (VR). Recent advances in generative AI have enabled the automatic creation of realistic 3D layouts. However, it is…

Human-Computer Interaction · Computer Science 2024-08-20 Lei Zhang , Jin Pan , Jacob Gettig , Steve Oney , Anhong Guo
‹ Prev 1 8 9 10 Next ›