English
Related papers

Related papers: A Transformer-based Multimodal Fusion Model for Ef…

200 papers

Recently, Transformer-based methods have shown impressive performance in single image super-resolution (SISR) tasks due to the ability of global feature extraction. However, the capabilities of Transformers that need to incorporate…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Wenjie Li , Juncheng Li , Guangwei Gao , Jiantao Zhou , Jian Yang , Guo-Jun Qi

How should representations from complementary sensors be integrated for autonomous driving? Geometry-based sensor fusion has shown great promise for perception tasks such as object detection and motion forecasting. However, for the actual…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Aditya Prakash , Kashyap Chitta , Andreas Geiger

Crowd counting is an important problem in computer vision due to its wide range of applications in image understanding. Currently, this problem is typically addressed using deep learning approaches, such as Convolutional Neural Networks…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Zhen Wang , Yuelei Li , Jia Wan , Nuno Vasconcelos

In this paper, a novel Unified Multi-Task Learning Framework of Real-Time Drone Supervision for Crowd Counting (MFCC) is proposed, which utilizes an image fusion network architecture to fuse images from the visible and thermal infrared…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Siqi Gu , Zhichao Lian

Crowd counting finds direct applications in real-world situations, making computational efficiency and performance crucial. However, most of the previous methods rely on a heavy backbone and a complex downstream architecture that restricts…

Computer Vision and Pattern Recognition · Computer Science 2024-01-12 Yashwardhan Chaudhuri , Ankit Kumar , Orchid Chetia Phukan , Arun Balaji Buduru

Multispectral image pairs can provide the combined information, making object detection applications more reliable and robust in the open world. To fully exploit the different modalities, we present a simple yet effective cross-modality…

Image and Video Processing · Electrical Eng. & Systems 2022-10-05 Fang Qingyun , Han Dapeng , Wang Zhaokui

This paper presents a hybrid model combining Transformer and CNN for predicting the current waveform in signal lines. Unlike traditional approaches such as current source models, driver linear representations, waveform functional fitting,…

Signal Processing · Electrical Eng. & Systems 2025-06-11 Junlang Huang , Hao Chen , Li Luo , Yong Cai , Lexin Zhang , Tianhao Ma , Yitian Zhang , Zhong Guan

RGB-Thermal (RGB-T) crowd counting is a challenging task, which uses thermal images as complementary information to RGB images to deal with the decreased performance of unimodal RGB-based methods in scenes with low-illumination or similar…

Computer Vision and Pattern Recognition · Computer Science 2022-08-16 Pengyu Chen , Junyu Gao , Yuan Yuan , Qi Wang

Image fusion is a technique to integrate information from multiple source images with complementary information to improve the richness of a single image. Due to insufficient task-specific training data and corresponding ground truth, most…

Computer Vision and Pattern Recognition · Computer Science 2022-01-20 Linhao Qu , Shaolei Liu , Manning Wang , Shiman Li , Siqi Yin , Qin Qiao , Zhijian Song

Automatic estimation of the number of people in unconstrained crowded scenes is a challenging task and one major difficulty stems from the huge scale variation of people. In this paper, we propose a novel Deep Structured Scale Integration…

Computer Vision and Pattern Recognition · Computer Science 2019-08-26 Lingbo Liu , Zhilin Qiu , Guanbin Li , Shufan Liu , Wanli Ouyang , Liang Lin

Recently, the research of wireless sensing has achieved more intelligent results, and the intelligent sensing of human location and activity can be realized by means of WiFi devices. However, most of the current human environment perception…

Machine Learning · Computer Science 2019-03-14 Shangqing Liu , Yanchao Zhao , Fanggang Xue , Bing Chen , Xiang Chen

Automatic segmentation of medical images based on multi-modality is an important topic for disease diagnosis. Although the convolutional neural network (CNN) has been proven to have excellent performance in image segmentation tasks, it is…

Computer Vision and Pattern Recognition · Computer Science 2022-04-27 Xuejian Li , Shiqiang Ma , Jijun Tang , Fei Guo

The classification of indoor scenes is a critical component in various applications, such as intelligent robotics for assistive living. While deep learning has significantly advanced this field, models often suffer from reduced performance…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Willams de Lima Costa , Raul Ismayilov , Nicola Strisciuglio , Estefania Talavera Martinez

This paper investigates the role of global context for crowd counting. Specifically, a pure transformer is used to extract features with global information from overlapping image patches. Inspired by classification, we add a context token…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Guolei Sun , Yun Liu , Thomas Probst , Danda Pani Paudel , Nikola Popovic , Luc Van Gool

Remote sensing image fusion aims to generate a high-resolution multi/hyper-spectral image by combining a high-resolution image with limited spectral data and a low-resolution image rich in spectral information. Current deep learning (DL)…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Siran Peng , Xiangyu Zhu , Haoyu Deng , Liang-Jian Deng , Zhen Lei

Crowd counting from unconstrained scene images is a crucial task in many real-world applications like urban surveillance and management, but it is greatly challenged by the camera's perspective that causes huge appearance variations in…

Computer Vision and Pattern Recognition · Computer Science 2018-07-03 Lingbo Liu , Hongjun Wang , Guanbin Li , Wanli Ouyang , Liang Lin

Recently multi-view crowd counting using deep neural networks has been proposed to enable counting in large and wide scenes using multiple cameras. The current methods project the camera-view features to the average-height plane of the 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Qi Zhang , Antoni B. Chan

In this paper, the dual-optical attention fusion crowd head point counting model (TAPNet) is proposed to address the problem of the difficulty of accurate counting in complex scenes such as crowd dense occlusion and low light in crowd…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Fei Zhou , Yi Li , Mingqing Zhu

This paper presents a novel approach for multimodal data fusion based on the Vector-Quantized Variational Autoencoder (VQVAE) architecture. The proposed method is simple yet effective in achieving excellent reconstruction performance on…

Machine Learning · Computer Science 2025-01-20 Mohammud J. Bocus , Xiaoyang Wang , Robert. J. Piechocki

Diffusion models have shown exceptional scaling properties in the image synthesis domain, and initial attempts have shown similar benefits for applying diffusion to unconditional text synthesis. Denoising diffusion models attempt to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-17 Matthew Baas , Kevin Eloff , Herman Kamper