English
Related papers

Related papers: Talk to Me, Not the Slides: A Real-Time Wearable A…

200 papers

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

Monitoring students' engagement and understanding their learning pace in a virtual classroom becomes challenging in the absence of direct eye contact between the students and the instructor. Continuous monitoring of eye gaze and gaze…

Human-Computer Interaction · Computer Science 2022-04-19 Snigdha Das , Sandip Chakraborty , Bivas Mitra

Eye gaze is considered an important indicator for understanding and predicting user behaviour, as well as directing their attention across various domains including advertisement design, human-computer interaction and film viewing. In this…

Software Engineering · Computer Science 2024-11-21 Karolina Trajkovska , Matjaž Kljun , Klen Čopič Pucihar

In video conferencing, human faces serve as the primary visual focal points, playing multifaceted roles that enhance visual communication and emotional connection. However, we argue that a human face is also a side channel, which can…

Cryptography and Security · Computer Science 2026-04-09 Yong Huang , Yanzhao Lu , Mingyang Chen , En Zhang , Jiazi Li , Wanqing Tu

Nonverbal communication, in particular eye contact, is a critical element of the music classroom, shown to keep students on task, coordinate musical flow, and communicate improvisational ideas. Unfortunately, this nonverbal aspect to…

Human-Computer Interaction · Computer Science 2021-05-24 Ross Greer , Shlomo Dubnov

We introduce and analyze a novel approach to the problem of speaker identification in multi-party recorded meetings. Given a speech segment and a set of available candidate profiles, we propose a novel data-driven way to model the distance…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-23 Nikolaos Flemotomos , Dimitrios Dimitriadis

Telepresence robots are used in various forms in various use-cases that helps to avoid physical human presence at the scene of action. In this work, we focus on a telepresence robot that can be used to attend a meeting remotely with a group…

Robotics · Computer Science 2020-06-30 Hrishav Bakul Barua , Chayan Sarkar , Achanna Anil Kumar , Arpan Pal , Balamuralidhar P

Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain limits their ability…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Alexandre Symeonidis-Herzig , Özge Mercanoğlu Sincan , Richard Bowden

Why do some speakers capture a room almost instantly while others fail to connect? The real-time architecture of audience engagement remains largely a black box. Here, we used motion-captured animations to present the pure nonverbal…

Human-Computer Interaction · Computer Science 2026-03-02 Ralf Schmälzle , Yuetong Du , Sue Lim , Gary Bente

Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic full-body motion based on text prompts, such as talking while…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Bohong Chen , Yumeng Li , Yao-Xiang Ding , Tianjia Shao , Kun Zhou

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Jiadong Liang , Feng Lu

Voice assistants have recently achieved remarkable commercial success. However, the current generation of these devices is typically capable of only reactive interactions. In other words, interactions have to be initiated by the user, which…

Engaging in smooth conversations with others is a crucial social skill. However, differences in knowledge between conversation participants can sometimes hinder effective communication. To tackle this issue, this study proposes a real-time…

Human-Computer Interaction · Computer Science 2025-06-23 Yuichiro Fujimoto

The goal of this work is to determine 'who spoke when' in real-world meetings. The method takes surround-view video and single or multi-channel audio as inputs, and generates robust diarisation outputs. To achieve this, we propose a novel…

Sound · Computer Science 2019-06-25 Joon Son Chung , Bong-Jin Lee , Icksang Han

Predicting when to initiate speech in real-world environments remains a fundamental challenge for conversational agents. We introduce EgoSpeak, a novel framework for real-time speech initiation prediction in egocentric streaming video. By…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Junhyeok Kim , Min Soo Kim , Jiwan Chung , Jungbin Cho , Jisoo Kim , Sungwoong Kim , Gyeongbo Sim , Youngjae Yu

Millions of people listen to podcasts, audio stories, and lectures, but editing speech remains tedious and time-consuming. Creators remove unnecessary words, cut tangential discussions, and even re-record speech to make recordings concise…

Human-Computer Interaction · Computer Science 2025-08-12 Karim Benharrak , Puyuan Peng , Amy Pavel

Current text-to-speech (TTS) models face a persistent limitation: autoregressive (AR) models suffer from low generation efficiency, while modern non-autoregressive (NAR) models experience high latency due to their unordered temporal nature.…

Sound · Computer Science 2026-03-17 Zhengyan Sheng , Zhihao Du , Shiliang Zhang , Zhijie Yan , Liping Chen

Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Jinting Wang , Jun Wang , Hei Victor Cheng , Li Liu

We address the problem of designing a conversational avatar capable of a sequence of casual conversations with older adults. Users at risk of loneliness, social anxiety or a sense of ennui may benefit from practicing such conversations in…

Artificial Intelligence · Computer Science 2019-06-27 S. Zahra Razavi , Lenhart K. Schubert , Benjamin Kane , Mohammad Rafayet Ali , Kimberly Van Orden , Tianyi Ma

Eye tracking is a key technology for gaze-based interactions in Extended Reality (XR), but traditional frame-based systems struggle to meet XR's demands for high accuracy, low latency, and power efficiency. Event cameras offer a promising…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Junyuan Ding , Ziteng Wang , Chang Gao , Min Liu , Qinyu Chen
‹ Prev 1 8 9 10 Next ›