中文
相关论文

相关论文: Comparison of a Head-Mounted Display and a Curved …

200 篇论文

Blind and low-vision (BLV) people use audio descriptions (ADs) to access videos. However, current ADs are unalterable by end users, thus are incapable of supporting BLV individuals' potentially diverse needs and preferences. This research…

人机交互 · 计算机科学 2024-08-22 Rosiana Natalie , Ruei-Che Chang , Smitha Sheshadri , Anhong Guo , Kotaro Hara

Immersive virtual reality (VR) emerges as a promising research and clinical tool. However, several studies suggest that VR induced adverse symptoms and effects (VRISE) may undermine the health and safety standards, and the reliability of…

人机交互 · 计算机科学 2021-01-21 Panagiotis Kourtesis , Simona Collina , Leonidas A. A. Doumas , Sarah E. MacPherson

Object selection is essential in virtual reality (VR) head-mounted displays (HMDs). Prior work mainly focuses on enhancing and evaluating techniques for selecting a single object in VR, leaving a gap in the techniques for multi-object…

人机交互 · 计算机科学 2024-10-23 Rongkai Shi , Yushi Wei , Xuning Hu , Yu Liu , Yong Yue , Lingyun Yu , Hai-Ning Liang

Auditory attention and selective phase-locking are central to human speech understanding in complex acoustic scenes and cocktail party settings, yet these capabilities in multilingual subjects remain poorly understood. While machine…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Sai Samrat Kankanala , Ram Chandra , Sriram Ganapathy

We conducted a large-scale study of human perceptual quality judgments of High Dynamic Range (HDR) and Standard Dynamic Range (SDR) videos subjected to scaling and compression levels and viewed on three different display devices. HDR videos…

图像与视频处理 · 电气工程与系统科学 2023-04-27 Joshua P. Ebenezer , Zaixi Shang , Yixu Chen , Yongjun Wu , Hai Wei , Sriram Sethuraman , Alan C. Bovik

Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual cues to infer or hallucinate acoustic information, rather than verifying the audio…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Xiaofei Wen , Wenjie Jacky Mo , Xingyu Fu , Rui Cai , Tinghui Zhu , Wendi Li , Yanan Xie , Muhao Chen , Peng Qi

Consumer 3D scanners and depth cameras are increasingly being used to generate content and avatars for Virtual Reality (VR) environments and avoid the inconveniences of hand modeling; however, it is sometimes difficult to evaluate…

人机交互 · 计算机科学 2017-02-01 Jacob Thorn , Rodrigo Pizarro , Bernhard Spanlang , Pablo Bermell-Garcia , Mar Gonzalez-Franco

Selection of occluded objects is a challenging problem in virtual reality, even more so if multiple objects are involved. With the advent of new artificial intelligence technologies, we explore the possibility of leveraging large language…

人机交互 · 计算机科学 2024-10-29 Junlong Chen , Jens Grubert , Per Ola Kristensson

While voice user interfaces offer increased accessibility due to hands-free and eyes-free interactions, older adults often have challenges such as constructing structured requests and perceiving how such devices operate. Voice-first user…

Multimodal audiovisual perception can enable new avenues for robotic manipulation, from better material classification to the imitation of demonstrations for which only audio signals are available (e.g., playing a tune by ear). However, to…

机器人学 · 计算机科学 2026-03-09 Luca Macesanu , Boueny Folefack , Samik Singh , Ruchira Ray , Ben Abbatematteo , Roberto Martín-Martín

This paper analyzes the joint assessment of quality, spatial and social presence, empathy, attitude, and attention in three conditions: (A)visualizing and rating the quality of contents in a Head-Mounted Display (HMD), (B)visualizing the…

多媒体 · 计算机科学 2022-02-10 Marta Orduna , Pablo Pérez , Jesús Gutiérrez , Narciso García

While videos have become increasingly prevalent in delivering information across different educational and professional contexts, individuals with ADHD often face attention challenges when watching informational videos due to the dynamic,…

人机交互 · 计算机科学 2025-11-25 Hanxiu 'Hazel' Zhu , Ruijia Chen , Yuhang Zhao

Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Liangbin Huang , Xiaohua Liao , Chaoqun Cui , Shijing Wang , Zhaolong Huang , Yanlong Du , Wenji Mao

The abundance and ease of utilizing sound, along with the fact that auditory clues reveal so much about what happens in the scene, make the audio-visual space a perfectly intuitive choice for self-supervised representation learning.…

计算机视觉与模式识别 · 计算机科学 2021-06-17 Mahdi M. Kalayeh , Nagendra Kamath , Lingyi Liu , Ashok Chandrashekar

When watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlying correlation between audio and visual events, which can be…

多媒体 · 计算机科学 2020-08-19 Ying Cheng , Ruize Wang , Zhihao Pan , Rui Feng , Yuejie Zhang

Understanding social interaction in video requires reasoning over a dynamic interplay of verbal and non-verbal cues: who is speaking, to whom, and with what gaze or gestures. While Multimodal Large Language Models (MLLMs) are natural…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Liangyang Ouyang , Yifei Huang , Mingfang Zhang , Caixin Kang , Ryosuke Furuta , Yoichi Sato

In this paper, we study the associations between human faces and voices. Audiovisual integration, specifically the integration of facial and vocal information is a well-researched area in neuroscience. It is shown that the overlapping…

计算机视觉与模式识别 · 计算机科学 2018-11-05 Changil Kim , Hijung Valentina Shin , Tae-Hyun Oh , Alexandre Kaspar , Mohamed Elgharib , Wojciech Matusik

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However, in a more realistic setting, when multiple faces are…

音频与语音处理 · 电气工程与系统科学 2022-05-12 Otavio Braga , Takaki Makino , Olivier Siohan , Hank Liao

Extended reality technology has become a useful tool in many applications, but still suffers from visual deviations that can hamper the utility of the technology. This paper discusses the types of persisting visual deviations experienced…

人机交互 · 计算机科学 2024-10-30 Rudy De-Xin de Lange , Roemer Martin Bien Bakker , Tanja Johanna Juliana Bos

Objective: Distorted loudness perception is one of the main complaints of hearing aid users. Being able to measure loudness perception correctly in the clinic is essential for fitting hearing aids. For this, experiments in the clinic should…

神经元与认知 · 定量生物学 2022-05-05 Gerard Llorach , Dirk Oetting , Matthias Vormann , Markus Meis , Volker Hohmann