English
Related papers

Related papers: Timestamp-independent Haptic-Visual Synchronizatio…

200 papers

Harmonization improves data consistency and is central to effective integration of diverse imaging data acquired across multiple sites. Recent deep learning techniques for harmonization are predominantly supervised in nature and hence…

Image and Video Processing · Electrical Eng. & Systems 2021-10-04 Siyuan Liu , Pew-Thian Yap

We introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene…

Humans process significantly more information through the sense of touch than through vision. Consequently, haptics for telemanipulation is poised to become essential in the coming years, as it offers operators an additional sensory channel…

Robotics · Computer Science 2025-02-25 Julien Mellet , Fabio Ruggiero , Vincenzo Lippiello

When humans perceive the world, they naturally integrate multiple audio-visual tasks within dynamic, real-world scenes. However, current works such as event localization, parsing, segmentation and question answering are mostly explored…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Guangyao Li , Xin Wang , Wenwu Zhu

In order to provide perceptually accurate multimodal feedback during driving situations, it is vital to understand the threshold at which drivers are able to recognize asyncrony between multiple incoming Stimuli. In this work, we…

Human-Computer Interaction · Computer Science 2023-07-12 Gyanendra Sharma , Hiroshi Yasuda , Manuel Kuehner

An important application of haptic technology to digital product development is in virtual prototyping (VP), part of which deals with interactive planning, simulation, and verification of assembly-related activities, collectively called…

Human-Computer Interaction · Computer Science 2017-12-12 Morad Behandish , Horea T. Ilies

A consistent spatial-temporal coordination across multiple agents is fundamental for collaborative perception, which seeks to improve perception abilities through information exchange among agents. To achieve this spatial-temporal…

Artificial Intelligence · Computer Science 2024-06-03 Zixing Lei , Zhenyang Ni , Ruize Han , Shuo Tang , Dingju Wang , Chen Feng , Siheng Chen , Yanfeng Wang

Haptic captioning is the task of generating natural language descriptions from haptic signals, such as vibrations, for use in virtual reality, accessibility, and rehabilitation applications. While previous multimodal research has focused…

Computation and Language · Computer Science 2026-01-15 Guimin Hu , Daniel Hershcovich , Hasti Seifi

From the patter of rain to the crunch of snow, the sounds we hear often convey the visual textures that appear within a scene. In this paper, we present a method for learning visual styles from unlabeled audio-visual data. Our model learns…

Computer Vision and Pattern Recognition · Computer Science 2022-05-11 Tingle Li , Yichen Liu , Andrew Owens , Hang Zhao

Machines that can represent and describe environmental soundscapes have practical potential, e.g., for audio tagging and captioning systems. Prevailing learning paradigms have been relying on parallel audio-text data, which is, however,…

Sound · Computer Science 2022-05-04 Yanpeng Zhao , Jack Hessel , Youngjae Yu , Ximing Lu , Rowan Zellers , Yejin Choi

A human operator using a manual control interface has ready access to their own command signal, both by efference copy and proprioception. In contrast, a human supervisor typically relies on visual information alone. We propose supplying a…

Human-Computer Interaction · Computer Science 2024-03-01 Alia Gilbert , Sachit Krishnan , R. Brent Gillespie

High-fidelity haptic feedback is essential for immersive virtual environments, yet authoring realistic tactile textures remains a significant bottleneck for designers. We introduce HapticMatch, a visual-to-tactile generation framework…

Human-Computer Interaction · Computer Science 2026-01-26 Mingxin Zhang , Yu Yao , Yasutoshi Makino , Hiroyuki Shinoda , Masashi Sugiyama

Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-02 George Sterpu , Christian Saam , Naomi Harte

Haptic sciences and technologies benefit greatly from comprehensive datasets that capture tactile stimuli under controlled, systematic conditions. However, existing haptic datasets collect data through uncontrolled exploration, which…

Human-Computer Interaction · Computer Science 2025-11-07 Michikuni Eguchi , Tomohiro Hayase , Yuichi Hiroi , Takefumi Hiraki

Synthesizing realistic human-object interaction motions is a critical problem in VR/AR and human animation. Unlike the commonly studied scenarios involving a single human or hand interacting with one object, we address a more generic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Wenkun He , Yun Liu , Ruitao Liu , Li Yi

The paramount challenge in audio-driven One-shot Talking Head Animation (ADOS-THA) lies in capturing subtle imperceptible changes between adjacent video frames. Inherently, the temporal relationship of adjacent audio clips is highly…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Zhihua Xu , Tianshui Chen , Zhijing Yang , Siyuan Peng , Keze Wang , Liang Lin

As a new communication paradigm, semantic communication has received widespread attention in communication fields. However, since the decoding of semantic signals relies on contextual knowledge, misalignment between the starting position of…

Signal Processing · Electrical Eng. & Systems 2023-12-19 Xiaoyi Liu , Haotai Liang , Chen Dong , Xiaodong Xu

Recent advancements in audio generation have enabled the creation of high-fidelity audio clips from free-form textual descriptions. However, temporal relationships, a critical feature for audio content, are currently underrepresented in…

Sound · Computer Science 2024-07-04 Zeyu Xie , Xuenan Xu , Zhizheng Wu , Mengyue Wu

Visual and audio signals often coexist in natural environments, forming audio-visual events (AVEs). Given a video, we aim to localize video segments containing an AVE and identify its category. In order to learn discriminative features for…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Jinxing Zhou , Liang Zheng , Yiran Zhong , Shijie Hao , Meng Wang

Current visual generation methods can produce high quality videos guided by texts. However, effectively controlling object dynamics remains a challenge. This work explores audio as a cue to generate temporally synchronized image animations.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Lin Zhang , Shentong Mo , Yijing Zhang , Pedro Morgado