English
Related papers

Related papers: Zoom and Shift are All You Need

200 papers

Classification using multimodal data arises in many machine learning applications. It is crucial not only to model cross-modal relationship effectively but also to ensure robustness against loss of part of data or modalities. In this paper,…

Machine Learning · Computer Science 2019-04-22 Jun-Ho Choi , Jong-Seok Lee

Multimodal image alignment is the process of finding spatial correspondences between images formed by different imaging techniques or under different conditions, to facilitate heterogeneous data fusion and correlative analysis. The…

Computer Vision and Pattern Recognition · Computer Science 2022-07-01 Johan Öfverstedt , Joakim Lindblad , Nataša Sladoje

Developing effective multimodal fusion approaches has become increasingly essential in many real-world scenarios, such as health care and finance. The key challenge is how to preserve the feature expressiveness in each modality while…

Machine Learning · Computer Science 2025-10-24 Tsai Hor Chan , Feng Wu , Yihang Chen , Guosheng Yin , Lequan Yu

This study introduces a pioneering methodology for human action recognition by harnessing deep neural network techniques and adaptive fusion strategies across multiple modalities, including RGB, optical flows, audio, and depth information.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Novanto Yudistira

Image fusion methods and metrics for their evaluation have conventionally used pixel-based or low-level features. However, for many applications, the aim of image fusion is to effectively combine the semantic content of the input images.…

Computer Vision and Pattern Recognition · Computer Science 2021-10-14 P. R. Hill , D. R. Bull

Multimedia collections are more than ever growing in size and diversity. Effective multimedia retrieval systems are thus critical to access these datasets from the end-user perspective and in a scalable way. We are interested in…

Information Retrieval · Computer Science 2014-01-28 Gabriela Csurka , Julien Ah-Pine , Stéphane Clinchant

Different modalities of medical images provide unique physiological and anatomical information for diseases. Multi-modal medical image fusion integrates useful information from different complementary medical images with different…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Yushen Xu , Xiaosong Li , Yuchun Wang , Xiaoqi Cheng , Huafeng Li , Haishu Tan

Visual object tracking, which is primarily based on visible light image sequences, encounters numerous challenges in complicated scenarios, such as low light conditions, high dynamic ranges, and background clutter. To address these…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Hongze Sun , Rui Liu , Wuque Cai , Jun Wang , Yue Wang , Huajin Tang , Yan Cui , Dezhong Yao , Daqing Guo

Multimodal sentiment analysis, a pivotal task in affective computing, seeks to understand human emotions by integrating cues from language, audio, and visual signals. While many recent approaches leverage complex attention mechanisms and…

Computation and Language · Computer Science 2025-05-09 Nischal Mandal , Yang Li

Multimodal learning integrates information from different modalities to enhance model performance, yet it often suffers from modality imbalance, where dominant modalities overshadow weaker ones during joint optimization. This paper reveals…

Machine Learning · Computer Science 2025-10-17 Xiaoyu Ma , Hao Chen

Multimodal fusion is crucial in joint decision-making systems for rendering holistic judgments. Since multimodal data changes in open environments, dynamic fusion has emerged and achieved remarkable progress in numerous applications.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Bing Cao , Yinan Xia , Yi Ding , Changqing Zhang , Qinghua Hu

Integrating information from multiple modalities is arguably one of the essential prerequisites for grounding artificial intelligence systems with an understanding of the real world. Recent advances in video transformers that jointly learn…

Computer Vision and Pattern Recognition · Computer Science 2023-11-15 Dota Tianai Dong , Mariya Toneva

A multimodal network encodes relationships between the same set of nodes in multiple settings, and network alignment is a powerful tool for transferring information and insight between a pair of networks. We propose a method for multimodal…

Social and Information Networks · Computer Science 2017-03-31 Huda Nassar , David F. Gleich

This study investigates a hybrid method for text classification that integrates deep feature extraction from large language models, multi-scale fusion through feature pyramids, and structured modeling with graph neural networks to enhance…

Computation and Language · Computer Science 2025-11-11 Xiangchen Song , Yulin Huang , Jinxu Guo , Yuchen Liu , Yaxuan Luan

The proliferation of artificial intelligence has enabled a diversity of applications that bridge the gap between digital and physical worlds. As physical environments are too complex to model through a single information acquisition…

Machine Learning · Computer Science 2025-08-11 Yu Zheng

Multimodal learning integrates data from diverse sensors to effectively harness information from different modalities. However, recent studies reveal that joint learning often overfits certain modalities while neglecting others, leading to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Feng Yu , Xiangyu Wu , Yang Yang , Jianfeng Lu

Recent learning-based approaches have achieved impressive results in the field of single-shot camera localization. However, how best to fuse multiple modalities (e.g., image and depth) and to deal with degraded or missing input are less…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Kaichen Zhou , Changhao Chen , Bing Wang , Muhamad Risqi U. Saputra , Niki Trigoni , Andrew Markham

Inter-modal interaction plays an indispensable role in multimodal sentiment analysis. Due to different modalities sequences are usually non-alignment, how to integrate relevant information of each modality to learn fusion representations…

Computation and Language · Computer Science 2022-12-23 Kaicheng Yang , Ruxuan Zhang , Hua Xu , Kai Gao

Multimodal sentiment analysis in videos is a key task in many real-world applications, which usually requires integrating multimodal streams including visual, verbal and acoustic behaviors. To improve the robustness of multimodal fusion,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Lianyang Ma , Yu Yao , Tao Liang , Tongliang Liu

In asymmetric retrieval systems, models with different capacities are deployed on platforms with different computational and storage resources. Despite the great progress, existing approaches still suffer from a dilemma between retrieval…

Image and Video Processing · Electrical Eng. & Systems 2024-03-04 Hui Wu , Min Wang , Wengang Zhou , Zhenbo Lu , Houqiang Li
‹ Prev 1 8 9 10 Next ›