English
Related papers

Related papers: Multimodal Visual-Tactile Representation Learning …

200 papers

Vision-language models pre-trained on large scale of unlabeled biomedical images and associated reports learn generalizable semantic representations. These multi-modal representations can benefit various downstream tasks in the biomedical…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Xinliu Zhong , Kayhan Batmanghelich , Li Sun

Having access to multi-modal cues (e.g. vision and audio) empowers some cognitive tasks to be done faster compared to learning from a single modality. In this work, we propose to transfer knowledge across heterogeneous modalities, even…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Yanbei Chen , Yongqin Xian , A. Sophia Koepke , Ying Shan , Zeynep Akata

Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks. However, the heterogeneous nature of these modalities makes fusion challenging. Existing methods propose strategies to obtain…

Robotics · Computer Science 2025-07-22 Jinzhou Li , Tianhao Wu , Jiyao Zhang , Zeyuan Chen , Haotian Jin , Mingdong Wu , Yujun Shen , Yaodong Yang , Hao Dong

We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input. Existing visuo-tactile methods rely on global alignment and thus fail to capture the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Seongyu Kim , Seungwoo Lee , Hyeonggon Ryu , Joon Son Chung , Arda Senocak

Compared to rigid hands, underactuated compliant hands offer greater adaptability to object shapes, provide stable grasps, and are often more cost-effective. However, they introduce uncertainties in hand-object interactions due to their…

Robotics · Computer Science 2025-03-04 Osher Azulay , Dhruv Metha Ramesh , Nimrod Curtis , Avishai Sintov

Unsupervised representation learning methods like SwAV are proved to be effective in learning visual semantics of a target dataset. The main idea behind these methods is that different views of a same image represent the same semantics. In…

Computer Vision and Pattern Recognition · Computer Science 2022-06-13 Mehdi Seyfi , Amin Banitalebi-Dehkordi , Yong Zhang

Though the success of CLIP-based training recipes in vision-language models, their scalability to more modalities (e.g., 3D, audio, etc.) is limited to large-scale data, which is expensive or even inapplicable for rare modalities. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Weixian Lei , Yixiao Ge , Jianfeng Zhang , Dylan Sun , Kun Yi , Ying Shan , Mike Zheng Shou

Recently, many multi-modal trackers prioritize RGB as the dominant modality, treating other modalities as auxiliary, and fine-tuning separately various multi-modal tasks. This imbalance in modality dependence limits the ability of methods…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Xiantao Hu , Bineng Zhong , Qihua Liang , Zhiyi Mo , Liangtao Shi , Ying Tai , Jian Yang

We propose ViC-MAE, a model that combines both Masked AutoEncoders (MAE) and contrastive learning. ViC-MAE is trained using a global featured obtained by pooling the local representations learned under an MAE reconstruction loss and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Jefferson Hernandez , Ruben Villegas , Vicente Ordonez

Visual-tactile learning (VTL) enables embodied agents to perceive the physical world by integrating visual (VIS) and tactile (TAC) sensors. However, VTL still suffers from modality discrepancies between VIS and TAC images, as well as domain…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Liuxiang Qiu , Hui Da , Yuzhen Niu , Tiesong Zhao , Yang Cao , Zheng-Jun Zha

With the advent of large-scale multimodal video datasets, especially sequences with audio or transcribed speech, there has been a growing interest in self-supervised learning of video representations. Most prior work formulates the…

Computer Vision and Pattern Recognition · Computer Science 2020-09-21 Bruno Korbar , Fabio Petroni , Rohit Girdhar , Lorenzo Torresani

In this paper we present a self-supervised method for representation learning utilizing two different modalities. Based on the observation that cross-modal information has a high semantic meaning we propose a method to effectively exploit…

Computer Vision and Pattern Recognition · Computer Science 2019-04-30 Nawid Sayed , Biagio Brattoli , Björn Ommer

For successful deployment of robots in multifaceted situations, an understanding of the robot for its environment is indispensable. With advancing performance of state-of-the-art object detectors, the capability of robots to detect objects…

Human-Computer Interaction · Computer Science 2023-03-02 Daniel Weber , Wolfgang Fuhl , Enkelejda Kasneci , Andreas Zell

Perception is essential for the active interaction of physical agents with the external environment. The integration of multiple sensory modalities, such as touch and vision, enhances this perceptual process, creating a more comprehensive…

Robotics · Computer Science 2025-02-10 Enrico Donato , Egidio Falotico , Thomas George Thuruthel

We propose a self-supervised approach for learning representations and robotic behaviors entirely from unlabeled videos recorded from multiple viewpoints, and study how this representation can be used in two robotic imitation settings:…

Computer Vision and Pattern Recognition · Computer Science 2018-03-21 Pierre Sermanet , Corey Lynch , Yevgen Chebotar , Jasmine Hsu , Eric Jang , Stefan Schaal , Sergey Levine

Contact-rich manipulation tasks in unstructured environments often require both haptic and visual feedback. It is non-trivial to manually design a robot controller that combines these modalities which have very different characteristics.…

Contact-rich tasks continue to present many challenges for robotic manipulation. In this work, we leverage a multimodal visuotactile sensor within the framework of imitation learning (IL) to perform contact-rich tasks that involve relative…

Gaze following aims to interpret human-scene interactions by predicting the person's focal point of gaze. Prevailing approaches often adopt a two-stage framework, whereby multi-modality information is extracted in the initial stage for gaze…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Yuehao Song , Xinggang Wang , Jingfeng Yao , Wenyu Liu , Jinglin Zhang , Xiangmin Xu

In this paper, we propose Multi-View Dreaming, a novel reinforcement learning agent for integrated recognition and control from multi-view observations by extending Dreaming. Most current reinforcement learning method assumes a single-view…

Artificial Intelligence · Computer Science 2022-03-22 Akira Kinose , Masashi Okada , Ryo Okumura , Tadahiro Taniguchi

Video captioning aims to describe video contents using natural language format that involves understanding and interpreting scenes, actions and events that occurs simultaneously on the view. Current approaches have mainly concentrated on…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Antoine Hanna-Asaad , Decky Aspandi , Titus Zaharia
‹ Prev 1 3 4 5 6 7 10 Next ›