English
Related papers

Related papers: GMF: General Multimodal Fusion Framework for Corre…

200 papers

Image fusion technology is widely used to fuse the complementary information between multi-source remote sensing images. Inspired by the frontier of deep learning, this paper first proposes a heterogeneous-integrated framework based on a…

Image and Video Processing · Electrical Eng. & Systems 2024-05-15 Menghui Jiang , Huanfeng Shen , Jie Li , Liangpei Zhang

Point cloud processing is a challenging task due to its sparsity and irregularity. Prior works introduce delicate designs on either local feature aggregator or global geometric architecture, but few combine both advantages. We propose…

Computer Vision and Pattern Recognition · Computer Science 2022-05-17 Renrui Zhang , Ziyao Zeng , Ziyu Guo , Xinben Gao , Kexue Fu , Jianbo Shi

Point cloud anomaly detection is essential for various industrial applications. The huge computation and storage costs caused by the increasing product classes limit the application of single-class unsupervised methods, necessitating the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Yuqi Cheng , Yunkang Cao , Dongfang Wang , Weiming Shen , Wenlong Li

Infrared and visible image fusion has garnered considerable attention owing to the strong complementarity of these two modalities in complex, harsh environments. While deep learning-based fusion methods have made remarkable advances in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Guihui Li , Bowei Dong , Kaizhi Dong , Jiayi Li , Haiyong Zheng

Camouflage is a common visual phenomenon, which refers to hiding the foreground objects into the background images, making them briefly invisible to the human eye. Previous work has typically been implemented by an iterative optimization…

Computer Vision and Pattern Recognition · Computer Science 2022-03-21 Yangyang Li , Wei Zhai , Yang Cao , Zheng-jun Zha

We propose a compact and effective framework to fuse multimodal features at multiple layers in a single network. The framework consists of two innovative fusion schemes. Firstly, unlike existing multimodal methods that necessitate…

Computer Vision and Pattern Recognition · Computer Science 2021-08-12 Yikai Wang , Fuchun Sun , Ming Lu , Anbang Yao

Multi-modal fusion methods often suffer from two types of representation collapse: feature collapse where individual dimensions lose their discriminative power (as measured by eigenspectra), and modality collapse where one dominant modality…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Seulgi Kim , Kiran Kokilepersaud , Mohit Prabhushankar , Ghassan AlRegib

Modern industrial recommendation systems improve recommendation performance by integrating multimodal representations from pre-trained models into ID-based Click-Through Rate (CTR) prediction frameworks. However, existing approaches…

Information Retrieval · Computer Science 2026-04-17 Alin Fan , Hanqing Li , Sihan Lu , Jingsong Yuan , Jiandong Zhang

Recent advances in the industrial inspection of textured surfaces-in the form of visual inspection-have made such inspections possible for efficient, flexible manufacturing systems. We propose an unsupervised feature memory rearrangement…

Computer Vision and Pattern Recognition · Computer Science 2022-06-23 Haiming Yao , Wenyong Yu , Xue Wang

Representation learning of textual networks poses a significant challenge as it involves capturing amalgamated information from two modalities: (i) underlying network structure, and (ii) node textual attributes. For this, most existing…

Computation and Language · Computer Science 2020-11-06 Tony Gracious , Ambedkar Dukkipati

Transformers are a popular choice for classification tasks and as backbones for object detection tasks. However, their high latency brings challenges in their adaptation to lightweight object detection systems. We present an approximation…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Dharma KC , Venkata Ravi Kiran Dayana , Meng-Lin Wu , Venkateswara Rao Cherukuri , Hau Hwang

Recent advances in self-attention and pure multi-layer perceptrons (MLP) models for vision have shown great potential in achieving promising performance with fewer inductive biases. These models are generally based on learning interaction…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Yongming Rao , Wenliang Zhao , Zheng Zhu , Jiwen Lu , Jie Zhou

3D Gaussian Splatting (3DGS) is a powerful alternative to Neural Radiance Fields (NeRF), excelling in complex scene reconstruction and efficient rendering. However, it relies on high-quality point clouds from Structure-from-Motion (SfM),…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Ziao Liu , Zhenjia Li , Yifeng Shi , Xiangang Li

Effective and efficient task planning is essential for mobile robots, especially in applications like warehouse retrieval and environmental monitoring. These tasks often involve selecting one location from each of several target clusters,…

Artificial Intelligence · Computer Science 2026-03-23 Jiaqi Cheng , Mingfeng Fan , Xuefeng Zhang , Jingsong Liang , Yuhong Cao , Guohua Wu , Guillaume Adrien Sartoretti

Deepfake detection is crucial for curbing the harm it causes to society. However, current Deepfake detection methods fail to thoroughly explore artifact information across different domains due to insufficient intrinsic interactions. These…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Xueqi Qiu , Xingyu Miao , Fan Wan , Haoran Duan , Tejal Shah , Varun Ojhab , Yang Longa , Rajiv Ranjan

Traditional and deep learning-based fusion methods generated the intermediate decision map to obtain the fusion image through a series of post-processing procedures. However, the fusion results generated by these methods are easy to lose…

Computer Vision and Pattern Recognition · Computer Science 2021-04-21 Yongsheng Zang , Dongming Zhou , Changcheng Wang , Rencan Nie , Yanbu Guo

Multimodal large language models (MLLMs) typically rely on a single late-layer feature from a frozen vision encoder, leaving the encoder's rich hierarchy of visual cues under-utilized. MLLMs still suffer from visually ungrounded…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Chenchen Lin , Sanbao Su , Rachel Luo , Yuxiao Chen , Yan Wang , Marco Pavone , Fei Miao

Real-time open-vocabulary scene understanding is essential for efficient 3D perception in applications such as vision-language navigation, embodied intelligence, and augmented reality. However, existing methods suffer from imprecise…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Xiaofeng Jin , Matteo Frosi , Matteo Matteucci

We present a novel strategy for detecting global outliers in a federated learning setting, targeting in particular cross-silo scenarios. Our approach involves the use of two servers and the transmission of masked local data from clients to…

Machine Learning · Computer Science 2024-09-23 Daniele Malpetti , Laura Azzimonti

Point cloud registration methods can effectively handle large-scale, partially overlapping point cloud pairs. Despite its practicality, matching the unbalanced pairs in terms of spatial extent and density has been overlooked and rarely…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Kanghee Lee , Junha Lee , Jaesik Park