English
Related papers

Related papers: MIFNet: Learning Modality-Invariant Features for G…

200 papers

State-of-the-art deep learning algorithms yield remarkable results in many visual recognition tasks. However, they still fail to provide satisfactory results in scarce data regimes. To a certain extent this lack of data can be compensated…

Computer Vision and Pattern Recognition · Computer Science 2018-11-26 Frederik Pahde , Oleksiy Ostapenko , Patrick Jähnichen , Tassilo Klein , Moin Nabi

This work proposes a new end-to-end DCNN based approach for motion segmentation, especially for video sequences captured with such non-static cameras, called MOSNET. While other approaches focus on spatial or temporal context only, the…

Computer Vision and Pattern Recognition · Computer Science 2021-02-23 Markus Bosch

In recent years, predicting Big Five personality traits from multimodal data has received significant attention in artificial intelligence (AI). However, existing computational models often fail to achieve satisfactory performance.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Bin Tang , Keqi Pan , Miao Zheng , Ning Zhou , Jialu Sui , Dandan Zhu , Cheng-Long Deng , Shu-Guang Kuai

We propose a compact and effective framework to fuse multimodal features at multiple layers in a single network. The framework consists of two innovative fusion schemes. Firstly, unlike existing multimodal methods that necessitate…

Computer Vision and Pattern Recognition · Computer Science 2021-08-12 Yikai Wang , Fuchun Sun , Ming Lu , Anbang Yao

Visible-Infrared person re-identification (VI-ReID) is an important and challenging task in intelligent video surveillance. Existing methods mainly focus on learning a shared feature space to reduce the modality discrepancy between visible…

Computer Vision and Pattern Recognition · Computer Science 2023-08-30 Haichao Shi , Mandi Luo , Xiao-Yu Zhang , Ran He

Non-rigid inter-modality registration can facilitate accurate information fusion from different modalities, but it is challenging due to the very different image appearances across modalities. In this paper, we propose to train a non-rigid…

Computer Vision and Pattern Recognition · Computer Science 2018-05-01 Xiaohuan Cao , Jianhua Yang , Li Wang , Zhong Xue , Qian Wang , Dinggang Shen

The perception system for autonomous driving generally requires to handle multiple diverse sub-tasks. However, current algorithms typically tackle individual sub-tasks separately, which leads to low efficiency when aiming at obtaining…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Xuesong Chen , Shaoshuai Shi , Tao Ma , Jingqiu Zhou , Simon See , Ka Chun Cheung , Hongsheng Li

Despite the good results that have been achieved in unimodal segmentation, the inherent limitations of individual data increase the difficulty of achieving breakthroughs in performance. For that reason, multi-modal learning is increasingly…

Image and Video Processing · Electrical Eng. & Systems 2024-04-16 Yameng Wang , Yi Wan , Yongjun Zhang , Bin Zhang , Zhi Gao

Multi-Object Tracking (MOT) remains a vital component of intelligent video analysis, which aims to locate targets and maintain a consistent identity for each target throughout a video sequence. Existing works usually learn a discriminative…

Computer Vision and Pattern Recognition · Computer Science 2023-11-20 Yizhe Li , Sanping Zhou , Zheng Qin , Le Wang , Jinjun Wang , Nanning Zheng

Technological advances in medical data collection, such as high-throughput genomic sequencing and digital high-resolution histopathology, have contributed to the rising requirement for multimodal biomedical modelling, specifically for…

Machine Learning · Computer Science 2024-10-29 Konstantin Hemker , Nikola Simidjievski , Mateja Jamnik

Accurate multispectral image matching presents significant challenges due to non-linear intensity variations across spectral modalities, extreme viewpoint changes, and the scarcity of labeled datasets. Current state-of-the-art methods are…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Ismail Can Yagmur , Hasan F. Ates , Bahadir K. Gunturk

Deep Unfolding Network-based methods have emerged as effective solutions for multi-source image fusion by combining model-driven iterative optimization with data-driven deep learning. However, most existing deep unfolding image fusion…

Image and Video Processing · Electrical Eng. & Systems 2026-05-04 Ge Luo , Jun-Jie Huang , Qi Yu , Tianrui Liu , Ke Liang , Yuming Xiang , Wentao Zhao , Xinwang Liu , Meng Wang

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

Multimodal Machine Translation (MMT) has demonstrated the significant help of visual information in machine translation. However, existing MMT methods face challenges in leveraging the modality gap by enforcing rigid visual-linguistic…

Computation and Language · Computer Science 2025-10-09 Jiafeng Xiong , Yuting Zhao

In recent years, monocular depth estimation is applied to understand the surrounding 3D environment and has made great progress. However, there is an ill-posed problem on how to gain depth information directly from a single image. With the…

Computer Vision and Pattern Recognition · Computer Science 2021-07-15 Meiqi Pei

We study many-class few-shot (MCFS) problem in both supervised learning and meta-learning settings. Compared to the well-studied many-class many-shot and few-class few-shot problems, the MCFS problem commonly occurs in practical…

Machine Learning · Computer Science 2020-11-10 Lu Liu , Tianyi Zhou , Guodong Long , Jing Jiang , Chengqi Zhang

The focus of this survey is on the analysis of two modalities of multimodal deep learning: image and text. Unlike classic reviews of deep learning where monomodal image classifiers such as VGG, ResNet and Inception module are central…

Computer Vision and Pattern Recognition · Computer Science 2020-10-19 Wei Chen , Weiping Wang , Li Liu , Michael S. Lew

The integration of different imaging modalities, such as structural, diffusion tensor, and functional magnetic resonance imaging, with deep learning models has yielded promising outcomes in discerning phenotypic characteristics and…

Image and Video Processing · Electrical Eng. & Systems 2024-10-08 Zhiyuan Li , Hailong Li , Anca L. Ralescu , Jonathan R. Dillman , Mekibib Altaye , Kim M. Cecil , Nehal A. Parikh , Lili He

Obtaining common representations from different modalities is important in that they are interchangeable with each other in a classification problem. For example, we can train a classifier on image features in the common representations and…

Machine Learning · Computer Science 2016-12-30 Kuniaki Saito , Yusuke Mukuta , Yoshitaka Ushiku , Tatsuya Harada

Multimodal fake news detection has attracted many research interests in social forensics. Many existing approaches introduce tailored attention mechanisms to guide the fusion of unimodal features. However, how the similarity of these…

Computer Vision and Pattern Recognition · Computer Science 2022-05-31 Yangming Zhou , Qichao Ying , Zhenxing Qian , Sheng Li , Xinpeng Zhang
‹ Prev 1 3 4 5 6 7 10 Next ›