English
Related papers

Related papers: 52-Hz Whale Song: An Embodied VR Experience for Ex…

200 papers

Immersive audio-visual perception relies on the spatial integration of both auditory and visual information which are heterogeneous sensing modalities with different fields of reception and spatial resolution. This study investigates the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-17 Davide Berghi , Hanne Stenzel , Marco Volino , Adrian Hilton , Philip J. B. Jackson

Our experience of the world is multisensory, spanning a synthesis of language, sight, sound, touch, taste, and smell. Yet, artificial intelligence has primarily advanced in digital modalities like text, vision, and audio. This paper…

Machine Learning · Computer Science 2026-01-14 Paul Pu Liang

A visual metaphor constitutes a high-order form of human creativity, employing cross-domain semantic fusion to transform abstract concepts into impactful visual rhetoric. Despite the remarkable progress of generative AI, existing models…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Yu Xu , Yuxin Zhang , Juan Cao , Lin Gao , Chunyu Wang , Oliver Deussen , Tong-Yee Lee , Fan Tang

Social virtual reality (VR) serves as a vital platform for transgender individuals to explore their identities through avatars and foster personal connections within online communities. However, it presents a challenge: the disconnect…

Human-Computer Interaction · Computer Science 2024-02-14 Kassie Povinelli , Yuhang Zhao

As chatbots continue to evolve toward human-like, real-world, interactions, multimodality remains an active area of research and exploration. So far, efforts to integrate multimodality into chatbots have primarily focused on image-centric…

Computation and Language · Computer Science 2025-06-03 Jihyoung Jang , Minwook Bae , Minji Kim , Dilek Hakkani-Tur , Hyounghun Kim

Institutional and social barriers in higher education often prevent students with disabilities from effectively accessing support, including lengthy procedures, insufficient information, and high social-emotional demands. This study…

Human-Computer Interaction · Computer Science 2026-01-23 Alva Markelius , Fethiye Irmak Doğan , Julie Bailey , Guy Laban , Jenny L. Gibson , Hatice Gunes

The cross-speaker emotion transfer task in text-to-speech (TTS) synthesis particularly aims to synthesize speech for a target speaker with the emotion transferred from reference speech recorded by another (source) speaker. During the…

Sound · Computer Science 2022-04-11 Tao Li , Xinsheng Wang , Qicong Xie , Zhichao Wang , Lei Xie

As human-robot collaboration opportunities continue to expand, trust becomes ever more important for full engagement and utilization of robots. Affective trust, built on emotional relationship and interpersonal bonds is particularly…

Human-Computer Interaction · Computer Science 2020-01-17 Richard Savery , Ryan Rose , Gil Weinberg

Human acceptance of social robots is greatly effected by empathy and perceived understanding. This necessitates accurate and flexible responses to various input data from the user. While systems such as this can become increasingly complex…

Robotics · Computer Science 2024-12-31 Jordan Sinclair , Christopher Reardon

Voice style transfer, also called voice conversion, seeks to modify one speaker's voice to generate speech as if it came from another (target) speaker. Previous works have made progress on voice conversion with parallel training data and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-18 Siyang Yuan , Pengyu Cheng , Ruiyi Zhang , Weituo Hao , Zhe Gan , Lawrence Carin

Pseudo-haptics exploit carefully crafted visual or auditory cues to trick the brain into "feeling" forces that are never physically applied, offering a low-cost alternative to traditional haptic hardware. Here, we present a comparative…

Human-Computer Interaction · Computer Science 2025-10-13 Nishant Gautam , Somya Sharma , Peter Corcoran , Kaspar Althoefer

Confusion is a mental state triggered by cognitive disequilibrium that can occur in many types of task-oriented interaction, including Human-Robot Interaction (HRI). People may become confused while interacting with robots due to…

Computation and Language · Computer Science 2022-08-22 Na Li , Robert Ross

Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to…

Sound · Computer Science 2025-05-20 Yifan Hu , Rui Liu , Yi Ren , Xiang Yin , Haizhou Li

We report on an investigation into how different types of failures in a voice user interface (VUI) affects user frustration. To this end, we conducted a pilot user study ($n=10$) and a main user study ($n=30$), both with a simple…

Human-Computer Interaction · Computer Science 2020-02-11 Shiyoh Goetsu , Tetsuya Sakai

This study investigates the use of movement sensor data from a smart watch to infer an individual's emotional state. We present our findings on a user study with 50 participants. The experimental design is a mixed-design study;…

Human-Computer Interaction · Computer Science 2020-11-09 Juan C. Quiroz , Elena Geangu , Min Hooi Yong

Livestream shopping platforms often overlook the accessibility needs of the Deaf and Hard of Hearing (DHH) community, leading to barriers such as information inaccessibility and overload. To tackle these challenges, we developed…

Human-Computer Interaction · Computer Science 2025-08-12 Zeyu Yang , Zheng Wei , Yang Zhang , Xian Xu , Changyang He , Muzhi Zhou , Pan Hui

Applying existing methods to emotional support conversation -- which provides valuable assistance to people who are in need -- has two major limitations: (a) they generally employ a conversation-level emotion label, which is too…

Computation and Language · Computer Science 2022-04-01 Quan Tu , Yanran Li , Jianwei Cui , Bin Wang , Ji-Rong Wen , Rui Yan

Multimodal models play a key role in empathy detection, but their performance can suffer when modalities provide conflicting cues. To understand these failures, we examine cases where unimodal and multimodal predictions diverge. Using…

Computation and Language · Computer Science 2025-11-12 Maya Srikanth , Run Chen , Julia Hirschberg

Human emotions are complex and can be conveyed through nuanced touch gestures. Previous research has primarily focused on how humans recognize emotions through touch or on identifying key features of emotional expression for robots.…

Robotics · Computer Science 2025-08-13 Qiaoqiao Ren , Remko Proesmans , Yuanbo Hou , Francis wyffels , Tony Belpaeme

Large Language Models (LLMs) have substantially improved the conversational capabilities of social robots. Nevertheless, for an intuitive and fluent human-robot interaction, robots should be able to ground the conversation by relating…

Human-Computer Interaction · Computer Science 2026-04-09 Elisabeth Menendez , Michael Gienger , Santiago Martínez , Carlos Balaguer , Anna Belardinelli