中文
相关论文

相关论文: Comparison of a Head-Mounted Display and a Curved …

200 篇论文

High Dynamic Range (HDR) videos can represent a much greater range of brightness and color than Standard Dynamic Range (SDR) videos and are rapidly becoming an industry standard. HDR videos have more challenging capture, transmission, and…

图像与视频处理 · 电气工程与系统科学 2022-09-22 Zaixi Shang , Joshua P. Ebenezer , Alan C. Bovik , Yongjun Wu , Hai Wei , Sriram Sethuraman

Coordinated Multiple views (CMVs) are a visualization technique that simultaneously presents multiple visualizations in separate but linked views. There are many studies that report the advantages (e.g., usefulness for finding hidden…

人机交互 · 计算机科学 2022-04-21 Juyoung Oh , Chunggi Lee , Hwiyeon Kim , Kihwan Kim , Osang Kwon , Eric D. Ragan , Bum Chul Kwon , Sungahn Ko

Multimodal language analysis often considers relationships between features based on text and those based on acoustical and visual properties. Text features typically outperform non-text features in sentiment analysis or emotion recognition…

机器学习 · 计算机科学 2019-12-03 Zhongkai Sun , Prathusha Sarma , William Sethares , Yingyu Liang

This paper considers methods for audio display in a CAVE-type virtual reality theater, a 3 m cube with displays covering all six rigid faces. Headphones are possible since the user's headgear continuously measures ear positions, but…

声音 · 计算机科学 2011-06-08 Bowon Lee , Camille Goudeseune , Mark A. Hasegawa-Johnson

Does speaking style variation affect humans' ability to distinguish individuals from their voices? How do humans compare with automatic systems designed to discriminate between voices? In this paper, we attempt to answer these questions by…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Amber Afshan , Jody Kreiman , Abeer Alwan

Audio-visual speaker diarization aims at detecting "who spoke when" using both auditory and visual signals. Existing audio-visual diarization datasets are mainly focused on indoor environments like meeting rooms or news studios, which are…

计算机视觉与模式识别 · 计算机科学 2022-07-19 Eric Zhongcong Xu , Zeyang Song , Satoshi Tsutsui , Chao Feng , Mang Ye , Mike Zheng Shou

Student engagement is a key construct for learning and teaching. While most of the literature explored the student engagement analysis on computer-based settings, this paper extends that focus to classroom instruction. To best examine…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Ömer Sümer , Patricia Goldberg , Sidney D'Mello , Peter Gerjets , Ulrich Trautwein , Enkelejda Kasneci

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed…

音频与语音处理 · 电气工程与系统科学 2021-02-03 Hiroshi Sato , Tsubasa Ochiai , Keisuke Kinoshita , Marc Delcroix , Tomohiro Nakatani , Shoko Araki

Loudspeaker-based spatial audio reproduction schemes are increasingly used for evaluating hearing aids in complex acoustic conditions. To further establish the feasibility of this approach, this study investigated the interaction between…

声音 · 计算机科学 2015-08-04 Giso Grimm , Stephan Ewert , Volker Hohmann

Recognizing speech from silent lip movement, which is called lip reading, is a challenging task due to 1) the inherent information insufficiency of lip movement to fully represent the speech, and 2) the existence of homophenes that have…

计算机视觉与模式识别 · 计算机科学 2022-04-06 Minsu Kim , Jeong Hun Yeo , Yong Man Ro

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Qiushi Zhu , Jie Zhang , Yu Gu , Yuchen Hu , Lirong Dai

Students who are deaf and hard of hearing cannot hear in class and do not have full access to spoken information. They can use accommodations such as captions that display speech as text. However, compared with their hearing peers, the…

人机交互 · 计算机科学 2019-09-19 Raja Kushalnagar , Gary Behm , Kevin Wolfe , Peter Yeung , Becca Dingman , Shareef Ali , Abraham Glasser , Claire Ryan

We introduce and explore a new multimodal input representation for vision-language models: acoustic field video. Unlike conventional video (RGB with stereo/mono audio), our video stream provides a spatially grounded visualization of sound…

人机交互 · 计算机科学 2026-01-27 Daehwa Kim , Chris Harrison

Research in multi-modal interfaces aims to provide solutions to immersion and increase overall human performance. A promising direction is combining auditory, visual and haptic interaction between the user and the simulated environment.…

人机交互 · 计算机科学 2020-04-01 Eleftherios Triantafyllidis , Christopher McGreavy , Jiacheng Gu , Zhibin Li

We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual speech separation is a…

Predicting brain activity in response to naturalistic, multimodal stimuli is a key challenge in computational neuroscience. While encoding models are becoming more powerful, their ability to generalize to truly novel contexts remains a…

Audio-visual speech enhancement (AVSE) methods use both audio and visual features for the task of speech enhancement and the use of visual features has been shown to be particularly effective in multi-speaker scenarios. In the majority of…

音频与语音处理 · 电气工程与系统科学 2022-02-18 Shrishti Saha Shetu , Soumitro Chakrabarty , Emanuël A. P. Habets

The paper presents a pilot study conducted with spatial visual, audiovisual and auditory brain-computer-interface (BCI) based speller paradigms. The psychophysical experiments are conducted with healthy subjects in order to evaluate a…

人机交互 · 计算机科学 2012-10-11 Moonjeong Chang , Nozomu Nishikawa , Zhenyu Cai , Shoji Makino , Tomasz M. Rutkowski

Automatic Cued Speech Recognition (ACSR) provides an intelligent human-machine interface for visual communications, where the Cued Speech (CS) system utilizes lip movements and hand gestures to code spoken language for hearing-impaired…

计算机视觉与模式识别 · 计算机科学 2023-02-28 Lei Liu , Li Liu

We present Masked Audio-Video Learners (MAViL) to train audio-visual representations. Our approach learns with three complementary forms of self-supervision: (1) reconstruction of masked audio and video input data, (2) intra- and…

计算机视觉与模式识别 · 计算机科学 2023-07-18 Po-Yao Huang , Vasu Sharma , Hu Xu , Chaitanya Ryali , Haoqi Fan , Yanghao Li , Shang-Wen Li , Gargi Ghosh , Jitendra Malik , Christoph Feichtenhofer