English
Related papers

Related papers: Deep Latent Space Learning for Cross-modal Mapping…

200 papers

We present CrissCross, a self-supervised framework for learning audio-visual representations. A novel notion is introduced in our framework whereby in addition to learning the intra-modal and standard 'synchronous' cross-modal relations,…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Pritam Sarkar , Ali Etemad

We propose a transfer deep learning (TDL) framework that can transfer the knowledge obtained from a single-modal neural network to a network with a different modality. Specifically, we show that we can leverage speech data to fine-tune the…

Neural and Evolutionary Computing · Computer Science 2016-02-19 Seungwhan Moon , Suyoun Kim , Haohan Wang

Localization is an indispensable component of a robot's autonomy stack that enables it to determine where it is in the environment, essentially making it a precursor for any action execution or planning. Although convolutional neural…

Robotics · Computer Science 2018-03-13 Abhinav Valada , Noha Radwan , Wolfram Burgard

Audio-visual representation learning is an important task from the perspective of designing machines with the ability to understand complex events. To this end, we propose a novel multimodal framework that instantiates multiple instance…

Computer Vision and Pattern Recognition · Computer Science 2018-07-10 Sanjeel Parekh , Slim Essid , Alexey Ozerov , Ngoc Q. K. Duong , Patrick Pérez , Gaël Richard

Multi-modal learning relates information across observation modalities of the same physical phenomenon to leverage complementary information. Most multi-modal machine learning methods require that all the modalities used for training are…

Machine Learning · Computer Science 2021-03-10 Vandana Rajan , Alessio Brutti , Andrea Cavallaro

Heterogeneous data fusion can enhance the robustness and accuracy of an algorithm on a given task. However, due to the difference in various modalities, aligning the sensors and embedding their information into discriminative and compact…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Aditya Dutt , Alina Zare , Paul Gader

The deep learning-based speech enhancement (SE) methods always take the clean speech's waveform or time-frequency spectrum feature as the learning target, and train the deep neural network (DNN) by reducing the error loss between the DNN's…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-02 Yuewei Zhang , Huanbin Zou , Jie Zhu

Cross-modal learning has become a fundamental paradigm for integrating heterogeneous information sources such as images, text, and structured attributes. However, multimodal representations often suffer from modality dominance, redundant…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Xuecheng Li , Weikuan Jia , Alisher Kurbonaliev , Qurbonaliev Alisher , Khudzhamkulov Rustam , Ismoilov Shuhratjon , Eshmatov Javhariddin , Yuanjie Zheng

Hashing has been widely applied to multimodal retrieval on large-scale multimedia data due to its efficiency in computation and storage. In this article, we propose a novel deep semantic multimodal hashing network (DSMHN) for scalable…

Computer Vision and Pattern Recognition · Computer Science 2022-01-06 Lu Jin , Zechao Li , Jinhui Tang

Concurrent Speaker Detection (CSD), the task of identifying active speakers and their overlaps in an audio signal, is essential for various audio applications, including meeting transcription, speaker diarization, and speech separation.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-16 Amit Eliav , Sharon Gannot

This paper addresses the problem of semi-supervised transfer learning with limited cross-modality data in remote sensing. A large amount of multi-modal earth observation images, such as multispectral imagery (MSI) or synthetic aperture…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Danfeng Hong , Naoto Yokoya , Gui-Song Xia , Jocelyn Chanussot , Xiao Xiang Zhu

Transformer-based multimodal models are widely used in industrial-scale recommendation, search, and advertising systems for content understanding and relevance ranking. Enhancing labeled training data quality and cross-modal fusion…

Multimedia · Computer Science 2025-10-03 Yu Sun , Yin Li , Ruixiao Sun , Chunhui Liu , Fangming Zhou , Ze Jin , Linjie Wang , Xiang Shen , Zhuolin Hao , Hongyu Xiong

Multimodal deep learning systems which employ multiple modalities like text, image, audio, video, etc., are showing better performance in comparison with individual modalities (i.e., unimodal) systems. Multimodal machine learning involves…

Machine Learning · Computer Science 2022-01-19 Anil Rahate , Rahee Walambe , Sheela Ramanna , Ketan Kotecha

Recent deep learning models can efficiently combine inputs from different modalities (e.g., images and text) and learn to align their latent representations, or to translate signals from one domain to another (as in image captioning, or…

Artificial Intelligence · Computer Science 2025-11-27 Benjamin Devillers , Léopold Maytié , Rufin VanRullen

The Dynamic Saliency Prediction (DSP) task simulates the human selective attention mechanism to perceive the dynamic scene, which is significant and imperative in many vision tasks. Most of existing methods only consider visual cues, while…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Hailong Ning , Bin Zhao , Zhanxuan Hu , Lang He , Ercheng Pei

Visual navigation has received significant attention recently. Most of the prior works focus on predicting navigation actions based on semantic features extracted from visual encoders. However, these approaches often rely on large datasets…

Robotics · Computer Science 2024-03-19 Hongyu Li , Taskin Padir , Huaizu Jiang

The seen birds twitter, the running cars accompany with noise, etc. These naturally audiovisual correspondences provide the possibilities to explore and understand the outside world. However, the mixed multiple objects and sounds make it…

Computer Vision and Pattern Recognition · Computer Science 2019-04-22 Di Hu , Feiping Nie , Xuelong Li

We introduce a supervised-learning framework for non-rigid point set alignment of a new kind - Displacements on Voxels Networks (DispVoxNets) - which abstracts away from the point set representation and regresses 3D displacement fields on…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 Soshi Shimada , Vladislav Golyanik , Edgar Tretschk , Didier Stricker , Christian Theobalt

Recent advances in representation learning have demonstrated an ability to represent information from different modalities such as video, text, and audio in a single high-level embedding vector. In this work we present a self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2021-06-11 Alexander H. Liu , SouYoung Jin , Cheng-I Jeff Lai , Andrew Rouditchenko , Aude Oliva , James Glass

Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificially mixed video…

Computer Vision and Pattern Recognition · Computer Science 2019-08-22 Ruohan Gao , Kristen Grauman