中文
相关论文

相关论文: Hierarchical Audio-Visual-Proprioceptive Fusion fo…

200 篇论文

Contact-rich manipulation tasks in unstructured environments often require both haptic and visual feedback. However, it is non-trivial to manually design a robot controller that combines modalities with very different characteristics. While…

With the recently increasing capabilities of modern vehicles, novel approaches for interaction emerged that go beyond traditional touch-based and voice command approaches. Therefore, hand gestures, head pose, eye gaze, and speech have been…

人机交互 · 计算机科学 2022-11-08 Amr Gomaa

The research introduces a reproducible framework for transforming raw, heterogeneous sensor streams into aligned, semantically meaningful representations for multimodal human activity recognition. Grounded in the Carnegie Mellon University…

应用统计 · 统计学 2026-05-05 Yiyao Yang , Yasemin Gulbahar

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

Video caption refers to generating a descriptive sentence for a specific short video clip automatically, which has achieved remarkable success recently. However, most of the existing methods focus more on visual information while ignoring…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Wangli Hao , Zhaoxiang Zhang , He Guan , Guibo Zhu

Multi-modal learning is a fast growing area in artificial intelligence. It tries to help machines understand complex things by combining information from different sources, like images, text, and audio. By using the strengths of each…

Recent advances in robot imitation learning have yielded powerful visuomotor policies capable of manipulating a wide variety of objects directly from monocular visual inputs. However, monocular observations inherently lack reliable depth…

机器人学 · 计算机科学 2026-05-12 Evans Han , Yunfan Jiang , Yingke Wang , Haoyue Xiao , Huang Huang , Jianwen Xie , Jiajun Wu , Li Fei-Fei , Ruohan Zhang

This paper introduces a novel framework integrating nonlinear acoustic computing and reinforcement learning to enhance advanced human-robot interaction under complex noise and reverberation. Leveraging physically informed wave equations…

机器人学 · 计算机科学 2025-05-07 Xiaoliang Chen , Xin Yu , Le Chang , Yunhe Huang , Jiashuai He , Shibo Zhang , Jin Li , Likai Lin , Ziyu Zeng , Xianling Tu , Shuyu Zhang

As human-robot collaboration is becoming more widespread, there is a need for a more natural way of communicating with the robot. This includes combining data from several modalities together with the context of the situation and background…

人机交互 · 计算机科学 2024-04-03 Petr Vanc , Radoslav Skoviera , Karla Stepanova

Visuotactile sensing offers rich contact information that can help mitigate performance bottlenecks in imitation learning, particularly under vision-limited conditions, such as ambiguous visual cues or occlusions. Effectively fusing visual…

机器人学 · 计算机科学 2025-05-13 Shulong Jiang , Shiqi Zhao , Yuxuan Fan , Peng Yin

Multi-modal emotion recognition in conversations is a challenging problem due to the complex and complementary interactions between different modalities. Audio and textual cues are particularly important for understanding emotions from a…

声音 · 计算机科学 2025-04-02 Jiachen Luo , Huy Phan , Lin Wang , Joshua Reiss

Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks. However, the heterogeneous nature of these modalities makes fusion challenging. Existing methods propose strategies to obtain…

机器人学 · 计算机科学 2025-07-22 Jinzhou Li , Tianhao Wu , Jiyao Zhang , Zeyuan Chen , Haotian Jin , Mingdong Wu , Yujun Shen , Yaodong Yang , Hao Dong

Robust manipulation often hinges on a robot's ability to perceive extrinsic contacts-contacts between a grasped object and its surrounding environment. However, these contacts are difficult to observe through vision alone due to occlusions,…

机器人学 · 计算机科学 2025-10-01 Xili Yi , Jayjun Lee , Nima Fazeli

Existing top-performance autonomous driving systems typically rely on the multi-modal fusion strategy for reliable scene understanding. This design is however fundamentally restricted due to overlooking the modality-specific strengths and…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Zeyu Yang , Nan Song , Wei Li , Xiatian Zhu , Li Zhang , Philip H. S. Torr

Robotic manipulation tasks often rely on static cameras for perception, which can limit flexibility, particularly in scenarios like robotic surgery and cluttered environments where mounting static cameras is impractical. Ideally, robots…

机器人学 · 计算机科学 2025-09-18 Xiatao Sun , Francis Fan , Yinxing Chen , Daniel Rakita

Multi-modal fusion is a fundamental task for the perception of an autonomous driving system, which has recently intrigued many researchers. However, achieving a rather good performance is not an easy task due to the noisy raw data,…

计算机视觉与模式识别 · 计算机科学 2024-12-18 Keli Huang , Botian Shi , Xiang Li , Xin Li , Siyuan Huang , Yikang Li

Humans inherently possess generalizable visual representations that empower them to efficiently explore and interact with the environments in manipulation tasks. We advocate that such a representation automatically arises from…

机器人学 · 计算机科学 2023-10-05 Mingxiao Huo , Mingyu Ding , Chenfeng Xu , Thomas Tian , Xinghao Zhu , Yao Mu , Lingfeng Sun , Masayoshi Tomizuka , Wei Zhan

Multimodal emotion recognition has recently gained much attention since it can leverage diverse and complementary relationships over multiple modalities (e.g., audio, visual, biosignals, etc.), and can provide some robustness to noisy…

In robotics, Vision-Language-Action (VLA) models that integrate diverse multimodal signals from multi-view inputs have emerged as an effective approach. However, most prior work adopts static fusion that processes all visual inputs…

机器人学 · 计算机科学 2026-02-18 Young-Chae Son , Jung-Woo Lee , Yoon-Ji Choi , Dae-Kwan Ko , Soo-Chul Lim

Multimodal analysis has recently drawn much interest in affective computing, since it can improve the overall accuracy of emotion recognition over isolated uni-modal approaches. The most effective techniques for multimodal emotion…

计算机视觉与模式识别 · 计算机科学 2024-07-09 R. Gnana Praveen , Eric Granger , Patrick Cardinal