English
Related papers

Related papers: RGBT Tracking via Progressive Fusion Transformer w…

200 papers

This study introduces a pioneering methodology for human action recognition by harnessing deep neural network techniques and adaptive fusion strategies across multiple modalities, including RGB, optical flows, audio, and depth information.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Novanto Yudistira

Classifying the confusing samples in the course of RGBT tracking is a quite challenging problem, which hasn't got satisfied solution. Existing methods only focus on enlarging the boundary between positive and negative samples, however, the…

Computer Vision and Pattern Recognition · Computer Science 2020-03-18 Zhengzheng Tu , Chun Lin , Chenglong Li , Jin Tang , Bin Luo

The target representation learned by convolutional neural networks plays an important role in Thermal Infrared (TIR) tracking. Currently, most of the top-performing TIR trackers are still employing representations learned by the model…

Computer Vision and Pattern Recognition · Computer Science 2021-08-03 Jingxian Sun , Lichao Zhang , Yufei Zha , Abel Gonzalez-Garcia , Peng Zhang , Wei Huang , Yanning Zhang

Biometrics on mobile devices has attracted a lot of attention in recent years as it is considered a user-friendly authentication method. This interest has also been motivated by the success of Deep Learning (DL). Architectures based on…

Computer Vision and Pattern Recognition · Computer Science 2022-06-06 Paula Delgado-Santos , Ruben Tolosana , Richard Guest , Farzin Deravi , Ruben Vera-Rodriguez

Multi-modality image fusion aims at fusing modality-specific (complementarity) and modality-shared (correlation) information from multiple source images. To tackle the problem of the neglect of inter-feature relationships, high-frequency…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Xiaoli Zhang , Liying Wang , Libo Zhao , Xiongfei Li , Siwei Ma

The increasing adoption of human-robot interaction presents opportunities for technology to positively impact lives, particularly those with visual impairments, through applications such as guide-dog-like assistive robotics. We present a…

Robotics · Computer Science 2024-08-27 Adam Scicluna , Cedric Le Gentil , Sheila Sutjipto , Gavin Paul

Accurate traffic flow prediction is essential for applications like transport logistics but remains challenging due to complex spatio-temporal correlations and non-linear traffic patterns. Existing methods often model spatial and temporal…

Machine Learning · Computer Science 2025-03-18 Jing Chen , Haocheng Ye , Zhian Ying , Yuntao Sun , Wenqiang Xu

Accurate RGB-Thermal (RGB-T) crowd counting is crucial for public safety in challenging conditions. While recent Transformer-based methods excel at capturing global context, their inherent lack of spatial inductive bias causes attention to…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Yuhong Feng , Hongtao Chen , Qi Zhang , Jie Chen , Zhaoxi He , Mingzhe Liu , Jianghai Liao

Current RGBT tracking methods often overlook the impact of fusion location on mitigating modality gap, which is key factor to effective tracking. Our analysis reveals that shallower fusion yields smaller distribution gap. However, the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Andong Lu , Yuanzhi Guo , Wanyu Wang , Chenglong Li , Jin Tang , Bin Luo

Transformers have been successfully applied to the visual tracking task and significantly promote tracking performance. The self-attention mechanism designed to model long-range dependencies is the key to the success of Transformers.…

Computer Vision and Pattern Recognition · Computer Science 2022-05-10 Zhihong Fu , Zehua Fu , Qingjie Liu , Wenrui Cai , Yunhong Wang

Visual object tracking, which is primarily based on visible light image sequences, encounters numerous challenges in complicated scenarios, such as low light conditions, high dynamic ranges, and background clutter. To address these…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Hongze Sun , Rui Liu , Wuque Cai , Jun Wang , Yue Wang , Huajin Tang , Yan Cui , Dezhong Yao , Daqing Guo

We propose cross-modal attentive connections, a new dynamic and effective technique for multimodal representation learning from wearable data. Our solution can be integrated into any stage of the pipeline, i.e., after any convolutional…

Machine Learning · Computer Science 2022-06-10 Anubhav Bhatti , Behnam Behinaein , Paul Hungler , Ali Etemad

Audio-visual embodied navigation aims to enable an agent to autonomously localize and reach a sound source in unseen 3D environments by leveraging auditory cues. The key challenge of this task lies in effectively modeling the interaction…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Yi Wang , Yinfeng Yu , Bin Ren

We abstract the features (i.e. learned representations) of multi-modal data into 1) uni-modal features, which can be learned from uni-modal training, and 2) paired features, which can only be learned from cross-modal interactions.…

Computer Vision and Pattern Recognition · Computer Science 2023-06-26 Chenzhuang Du , Jiaye Teng , Tingle Li , Yichen Liu , Tianyuan Yuan , Yue Wang , Yang Yuan , Hang Zhao

Focus based methods have shown promising results for the task of depth estimation. However, most existing focus based depth estimation approaches depend on maximal sharpness of the focal stack. Out of focus information in the focal stack…

Computer Vision and Pattern Recognition · Computer Science 2021-04-14 Yongri Piao , Yukun Zhang , Miao Zhang , Xinxin Ji

Action recognition from multi-modal and multi-view observations holds significant potential for applications in surveillance, robotics, and smart environments. However, existing methods often fall short of addressing real-world challenges…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Trung Thanh Nguyen , Yasutomo Kawanishi , Vijay John , Takahiro Komamizu , Ichiro Ide

How to effectively fuse cross-modal information is the key problem for RGB-D salient object detection. Early fusion and the result fusion schemes fuse RGB and depth information at the input and output stages, respectively, hence incur the…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Nian Liu , Ni Zhang , Ling Shao , Junwei Han

Accurate beam prediction is essential for maintaining reliable links and high spectral efficiency in dynamic low-altitude wireless networks. However, existing approaches often fail to capture the deep correlations across heterogeneous…

Signal Processing · Electrical Eng. & Systems 2025-12-03 Xiaotong Zhao , Yuanhao Cui , Weijie Yuan , Ziye Jia , Heng Liu , Chengwen Xing

In the field of multimodal segmentation, the correlation between different modalities can be considered for improving the segmentation results. Considering the correlation between different MR modalities, in this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2021-11-10 Tongxue Zhou , Su Ruan , Pierre Vera , Stéphane Canu

Tasks that rely on multi-modal information typically include a fusion module that combines information from different modalities. In this work, we develop a Refiner Fusion Network (ReFNet) that enables fusion modules to combine strong…

Computer Vision and Pattern Recognition · Computer Science 2021-04-09 Sethuraman Sankaran , David Yang , Ser-Nam Lim