English
Related papers

Related papers: Talk to Me, Not the Slides: A Real-Time Wearable A…

200 papers

The think aloud method is an important and commonly used tool for usability optimization. However, analyzing think aloud data could be time consuming. In this paper, we put forth an automatic analysis of verbal protocols and test the link…

Human-Computer Interaction · Computer Science 2023-07-12 Supriya Murali , Tina Walber , Christoph Schaefer , Sezen Lim

We present RealityTalk, a system that augments real-time live presentations with speech-driven interactive virtual elements. Augmented presentations leverage embedded visuals and animation for engaging and expressive storytelling. However,…

Human-Computer Interaction · Computer Science 2022-08-15 Jian Liao , Adnan Karim , Shivesh Jadon , Rubaiat Habib Kazi , Ryo Suzuki

Wearable devices are transforming human capabilities by seamlessly augmenting cognitive functions. In this position paper, we propose a voice-based, interactive learning companion designed to amplify and extend cognitive abilities through…

Human-Computer Interaction · Computer Science 2025-04-25 Chitralekha Gupta , Hanjun Wu , Praveen Sasikumar , Shreyas Sridhar , Priambudi Bagaskara , Suranga Nanayakkara

In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling audio-visual synchronization cue, which depends on the video…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-13 Yidi Jiang , Ruijie Tao , Zexu Pan , Haizhou Li

While walking meetings offer a healthy alternative to sit-down meetings, they also pose practical challenges. Taking notes is difficult while walking, which limits the potential of walking meetings. To address this, we designed the Walking…

Human-Computer Interaction · Computer Science 2023-04-06 Luke Haliburton , Natalia Bartłomiejczyk , Paweł W. Woźniak , Albrecht Schmidt , Jasmin Niess

At least 360 million people worldwide have disabling hearing loss that frequently causes difficulties in day-to-day conversations. Hearing aids often fail to offer enough benefits and have low adoption rates. However, people with hearing…

Human-Computer Interaction · Computer Science 2018-06-05 Benjamin M. Gorman

Advanced multimodal AI agents can now collaborate with users to solve challenges in the world. Yet, these emerging contextual AI systems rely on explicit communication channels between the user and system. We hypothesize that implicit…

In this paper, we introduce EyeEcho, a minimally-obtrusive acoustic sensing system designed to enable glasses to continuously monitor facial expressions. It utilizes two pairs of speakers and microphones mounted on glasses, to emit encoded…

Human-Computer Interaction · Computer Science 2024-02-27 Ke Li , Ruidong Zhang , Siyuan Chen , Boao Chen , Mose Sakashita , François Guimbretière , Cheng Zhang

Eye movements provide a window into human behaviour, attention, and interaction dynamics. Challenges in real-world, multi-person environments have, however, restrained eye-tracking research predominantly to single-person, in-lab settings.…

Human-Computer Interaction · Computer Science 2025-05-27 Shreshth Saxena , Areez Visram , Neil Lobo , Zahid Mirza , Mehak Rafi Khan , Biranugan Pirabaharan , Alexander Nguyen , Lauren K. Fink

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

Sound · Computer Science 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee

In crowded settings, the human brain can focus on speech from a target speaker, given prior knowledge of how they sound. We introduce a novel intelligent hearable system that achieves this capability, enabling target speech hearing to…

Sound · Computer Science 2024-05-31 Bandhav Veluri , Malek Itani , Tuochao Chen , Takuya Yoshioka , Shyamnath Gollakota

What makes a talk successful? Is it the content or the presentation? We try to estimate the contribution of the speaker's oratory skills to the talk's success, while ignoring the content of the talk. By oratory skills we refer to facial…

Sound · Computer Science 2021-10-05 Tzvi Michelson , Shmuel Peleg

Equal access to digital technologies is critical for education, employment, and social participation. However, mainstream interfaces are visually oriented, creating steep learning curves and frequent obstacles for screen reader users, and…

Human-Computer Interaction · Computer Science 2026-01-27 Nan Chen , Jing Lu , Zilong Wang , Luna K. Qiu , Siming Chen , Yuqing Yang

Slide-based teaching is widely used in higher education, yet in online, hybrid, and asynchronous contexts, slides often lose instructor presence, narrative continuity, and expressive framing that help learners connect with course content.…

Human-Computer Interaction · Computer Science 2026-05-26 Xinxing Wu

Social interactions are a fundamental part of daily life and play a critical role in well-being. As emerging technologies offer opportunities to unobtrusively monitor behavior, there is growing interest in using them to better understand…

Human-Computer Interaction · Computer Science 2025-08-07 Md Sabbir Ahmed , Arafat Rahman , Mark Rucker , Laura E. Barnes

Current methods for active speak er detection focus on modeling short-term audiovisual information from a single speaker. Although this strategy can be enough for addressing single-speaker scenarios, it prevents accurate detection when the…

Computer Vision and Pattern Recognition · Computer Science 2020-05-21 Juan Leon Alcazar , Fabian Caba Heilbron , Long Mai , Federico Perazzi , Joon-Young Lee , Pablo Arbelaez , Bernard Ghanem

Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and…

Novice content creators often invest significant time recording expressive speech for social media videos. While recent advancements in text-to-speech (TTS) technology can generate highly realistic speech in various languages and accents,…

Human-Computer Interaction · Computer Science 2025-04-08 Stephen Brade , Sam Anderson , Rithesh Kumar , Zeyu Jin , Anh Truong

In this work, we present a novel audio-visual dataset for active speaker detection in the wild. A speaker is considered active when his or her face is visible and the voice is audible simultaneously. Although active speaker detection is a…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 You Jin Kim , Hee-Soo Heo , Soyeon Choe , Soo-Whan Chung , Yoohwan Kwon , Bong-Jin Lee , Youngki Kwon , Joon Son Chung

We propose a system displaying audience eye gaze and nod reactions for enhancing synchronous remote communication. Recently, we have had increasing opportunities to speak to others remotely. In contrast to offline situations, however,…

Human-Computer Interaction · Computer Science 2022-04-06 Kiyosu Maeda , Riku Arakawa , Jun Rekimoto