中文
相关论文

相关论文: SonoHaptics: An Audio-Haptic Cursor for Gaze-Based…

200 篇论文

In this study we describe a methodology to realize visual images cognition in the broader sense, by a cross-modal stimulation through the auditory channel. An original algorithm of conversion from bi-dimensional images to sounds has been…

神经元与认知 · 定量生物学 2017-05-16 Takahisa Kishino , Sun Zhe , Roberto Marchisio , Ruggero Micheletto

Tremendous progress in visual scene generation now turns a single image into an explorable 3D world, yet immersion remains incomplete without sound. We introduce Image2AVScene, the task of generating a 3D audio-visual scene from a single…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Derong Jin , Xiyi Chen , Ming C. Lin , Ruohan Gao

Multimodal large language models (LMMs) excel in world knowledge and problem-solving abilities. Through the use of a world-facing camera and contextual AI, emerging smart accessories aim to provide a seamless interface between humans and…

人机交互 · 计算机科学 2024-02-01 Robert Konrad , Nitish Padmanaban , J. Gabriel Buckmaster , Kevin C. Boyle , Gordon Wetzstein

Open-vocabulary panoptic reconstruction offers comprehensive scene understanding, enabling advances in embodied robotics and photorealistic simulation. In this paper, we propose PanopticRecon++, an end-to-end method that formulates panoptic…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Xuan Yu , Yuxuan Xie , Yili Liu , Haojian Lu , Rong Xiong , Yiyi Liao , Yue Wang

Visual speaker recognition based on lip motion offers a silent, hands-free, and behavior-driven biometric solution that remains effective even when acoustic cues are unavailable. Compared to traditional methods that rely heavily on…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Junguang Yao , Wenye Liu , Stjepan Picek , Yue Zheng

The contact-rich nature of manipulation makes it a significant challenge for robotic teleoperation. While haptic feedback is critical for contact-rich tasks, providing intuitive directional cues within wearable teleoperation interfaces…

机器人学 · 计算机科学 2026-04-01 Xiangshan Tan , Jingtian Ji , Tianchong Jiang , Pedro Lopes , Matthew R. Walter

Mapping and understanding complex 3D environments is fundamental to how autonomous systems perceive and interact with the physical world, requiring both precise geometric reconstruction and rich semantic comprehension. While existing 3D…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Naman Patel , Prashanth Krishnamurthy , Farshad Khorrami

Safe navigation for the visually impaired individuals remains a critical challenge, especially concerning head-level obstacles, which traditional mobility aids often fail to detect. We introduce GuideTouch, a compact, affordable, standalone…

Stereohaptic vibration is an innovative vibrotactile technology that extends the conventional tactile localization to the surrounding space, representing a virtual vibration source in the external environment. Previously, we have developed…

人机交互 · 计算机科学 2024-11-11 Gen Ohara , Masashi Konyo , Satoshi Tadokoro

When watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlying correlation between audio and visual events, which can be…

多媒体 · 计算机科学 2020-08-19 Ying Cheng , Ruize Wang , Zhihao Pan , Rui Feng , Yuejie Zhang

Audio-Visual Segmentation (AVS) aims to extract the sounding object from a video frame, which is represented by a pixel-wise segmentation mask for application scenarios such as multi-modal video editing, augmented reality, and intelligent…

图像与视频处理 · 电气工程与系统科学 2024-12-25 Zhaofeng Shi , Qingbo Wu , Fanman Meng , Linfeng Xu , Hongliang Li

We propose a new approach for interaction in Virtual Reality (VR) using mobile robots as proxies for haptic feedback. This approach allows VR users to have the experience of sharing and manipulating tangible physical objects with remote…

人机交互 · 计算机科学 2017-02-01 Zhenyi He , Fengyuan Zhu , Aaron Gaudette , Ken Perlin

While current personal smart devices excel in digital domains, they fall short in assisting users during human environment interaction. This paper proposes Heads Up eXperience (HUX), an AI system designed to bridge this gap, serving as a…

人机交互 · 计算机科学 2024-07-30 Sukanth K , Sudhiksha Kandavel Rajan , Rajashekhar V S , Gowdham Prabhakar

Virtual-reality (VR) and augmented-reality (AR) technology is increasingly combined with eye-tracking. This combination broadens both fields and opens up new areas of application, in which visual perception and related cognitive processes…

计算机视觉与模式识别 · 计算机科学 2021-03-22 Lena Stubbemann , Dominik Dürrschnabel , Robert Refflinghaus

The quantification of audio aesthetics remains a complex challenge in audio processing, primarily due to its subjective nature, which is influenced by human perception and cultural context. Traditional methods often depend on human…

We present a new object representation, called Dense RepPoints, that utilizes a large set of points to describe an object at multiple levels, including both box level and pixel level. Techniques are proposed to efficiently process these…

计算机视觉与模式识别 · 计算机科学 2020-05-19 Ze Yang , Yinghao Xu , Han Xue , Zheng Zhang , Raquel Urtasun , Liwei Wang , Stephen Lin , Han Hu

Combining 3D vision with tactile sensing could unlock a greater level of dexterity for robots and improve several manipulation tasks. However, obtaining a close-up 3D view of the location where manipulation contacts occur can be…

机器人学 · 计算机科学 2023-03-14 Etienne Roberge , Guillaume Fornes , Jean-Philippe Roberge

Sound plays a crucial role in enhancing user experience and immersiveness in Augmented Reality (AR). However, current platforms lack support for AR sound authoring due to limited interaction types, challenges in collecting and specifying…

人机交互 · 计算机科学 2024-08-13 Xia Su , Jon E. Froehlich , Eunyee Koh , Chang Xiao

In 3D user interfaces, reaching out to grab and manipulate something works great until it is out of reach. Indirect techniques like gaze and pinch offer an alternative for distant interaction, but do not provide the same immediacy or…

Grounding objects in images using visual cues is a well-established approach in computer vision, yet the potential of audio as a modality for object recognition and grounding remains underexplored. We introduce YOSS, "You Only Speak Once to…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Wenhao Yang , Jianguo Wei , Wenhuan Lu , Lei Li