English
Related papers

Related papers: Modality Mixer Exploiting Complementary Informatio…

200 papers

This study introduces a pioneering methodology for human action recognition by harnessing deep neural network techniques and adaptive fusion strategies across multiple modalities, including RGB, optical flows, audio, and depth information.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Novanto Yudistira

Moving Object Detection (MOD) is a critical vision task for successfully achieving safe autonomous driving. Despite plausible results of deep learning methods, most existing approaches are only frame-based and may fail to reach reasonable…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Zhuyun Zhou , Zongwei Wu , Rémi Boutteau , Fan Yang , Cédric Demonceaux , Dominique Ginhac

This paper presents a pure transformer-based approach, dubbed the Multi-Modal Video Transformer (MM-ViT), for video action recognition. Different from other schemes which solely utilize the decoded RGB frames, MM-ViT operates exclusively in…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Jiawei Chen , Chiu Man Ho

Multitemporal hyperspectral image unmixing (MTHU) holds significant importance in monitoring and analyzing the dynamic changes of surface. However, compared to single-temporal unmixing, the multitemporal approach demands comprehensive…

Image and Video Processing · Electrical Eng. & Systems 2024-07-16 Hang Li , Qiankun Dong , Xueshuo Xie , Xia Xu , Tao Li , Zhenwei Shi

Processing and fusing information among multi-modal is a very useful technique for achieving high performance in many computer vision problems. In order to tackle multi-modal information more effectively, we introduce a novel framework for…

Computer Vision and Pattern Recognition · Computer Science 2019-05-01 Dong Wang , Yuan Yuan , Qi Wang

In this paper, we study the cross-modal image retrieval, where the inputs contain a source image plus some text that describes certain modifications to this image and the desired image. Prior work usually uses a three-stage strategy to…

Computer Vision and Pattern Recognition · Computer Science 2021-03-11 Chunbin Gu , Jiajun Bu , Xixi Zhou , Chengwei Yao , Dongfang Ma , Zhi Yu , Xifeng Yan

With the continuous emergence of various social media platforms frequently used in daily life, the multimodal meme understanding (MMU) task has been garnering increasing attention. MMU aims to explore and comprehend the meanings of memes…

Computation and Language · Computer Science 2025-03-18 Li Zheng , Hao Fei , Ting Dai , Zuquan Peng , Fei Li , Huisheng Ma , Chong Teng , Donghong Ji

Multimodal information retrieval (MIR) faces inherent challenges due to the heterogeneity of data sources and the complexity of cross-modal alignment. While previous studies have identified modal gaps in feature spaces, a systematic…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Fanheng Kong , Jingyuan Zhang , Yahui Liu , Hongzhi Zhang , Shi Feng , Xiaocui Yang , Daling Wang , Yu Tian , Victoria W. , Fuzheng Zhang , Guorui Zhou

Cross-modal retrieval has become popular in recent years, particularly with the rise of multimedia. Generally, the information from each modality exhibits distinct representations and semantic information, which makes feature tends to be in…

Information Retrieval · Computer Science 2023-08-29 Zichen Yuan , Qi Shen , Bingyi Zheng , Yuting Liu , Linying Jiang , Guibing Guo

AI-synthesized text and images have gained significant attention, particularly due to the widespread dissemination of multi-modal manipulations on the internet, which has resulted in numerous negative impacts on society. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2024-10-28 Jiazhen Wang , Bin Liu , Changtao Miao , Zhiwei Zhao , Wanyi Zhuang , Qi Chu , Nenghai Yu

Radiologists must utilize multiple modal images for tumor segmentation and diagnosis due to the limitations of medical imaging and the diversity of tumor signals. This leads to the development of multimodal learning in segmentation.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Chuyun Shen , Wenhao Li , Haoqing Chen , Xiaoling Wang , Fengping Zhu , Yuxin Li , Xiangfeng Wang , Bo Jin

A novel deep neural network training paradigm that exploits the conjoint information in multiple heterogeneous sources is proposed. Specifically, in a RGB-D based action recognition task, it cooperatively trains a single convolutional…

Computer Vision and Pattern Recognition · Computer Science 2018-01-04 Pichao Wang , Wanqing Li , Jun Wan , Philip Ogunbona , Xinwang Liu

In this paper, we present Fusion-GCN, an approach for multimodal action recognition using Graph Convolutional Networks (GCNs). Action recognition methods based around GCNs recently yielded state-of-the-art performance for skeleton-based…

Computer Vision and Pattern Recognition · Computer Science 2021-09-28 Michael Duhme , Raphael Memmesheimer , Dietrich Paulus

Emotions play a crucial role in human behavior and decision-making, making emotion recognition a key area of interest in human-computer interaction (HCI). This study addresses the challenges of emotion recognition by integrating facial…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Zaitian Wang , Jian He , Yu Liang , Xiyuan Hu , Tianhao Peng , Kaixin Wang , Jiakai Wang , Chenlong Zhang , Weili Zhang , Shuang Niu , Xiaoyang Xie

U-Net and its extensions have achieved great success in medical image segmentation. However, due to the inherent local characteristics of ordinary convolution operations, U-Net encoder cannot effectively extract global context information.…

Image and Video Processing · Electrical Eng. & Systems 2024-10-28 Fenghe Tang , Lingtao Wang , Chunping Ning , Min Xian , Jianrui Ding

Visible-infrared person re-identification (VI-ReID) aims to match individuals across different camera modalities, a critical task in modern surveillance systems. While current VI-ReID methods focus on cross-modality matching, real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Mahdi Alehdaghi , Rajarshi Bhattacharya , Pourya Shamsolmoali , Rafael M. O. Cruz , Eric Granger

The use of multimodal data in assisted diagnosis and segmentation has emerged as a prominent area of interest in current research. However, one of the primary challenges is how to effectively fuse multimodal features. Most of the current…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Xinxin Fan , Lin Liu , Haoran Zhang

With the development of depth sensors in recent years, RGBD object tracking has received significant attention. Compared with the traditional RGB object tracking, the addition of the depth modality can effectively solve the target and…

Computer Vision and Pattern Recognition · Computer Science 2022-11-16 Shang Gao , Jinyu Yang , Zhe Li , Feng Zheng , Aleš Leonardis , Jingkuan Song

Efficient modal feature fusion strategy is the key to achieve accurate segmentation of brain glioma. However, due to the specificity of different MRI modes, it is difficult to carry out cross-modal fusion with large differences in modal…

Image and Video Processing · Electrical Eng. & Systems 2025-03-21 Dong Chen , Boyue Zhao , Yi Zhang , Meng Zhao

Multimodal sentiment analysis has a wide range of applications due to its information complementarity in multimodal interactions. Previous works focus more on investigating efficient joint representations, but they rarely consider the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-31 Rongfei Chen , Wenju Zhou , Yang Li , Huiyu Zhou
‹ Prev 1 4 5 6 7 8 10 Next ›