English
Related papers

Related papers: Transcending Fusion: A Multi-Scale Alignment Metho…

200 papers

Multimodal Sentiment Analysis (MSA) integrates complementary features from text, video, and audio for robust emotion understanding in human interactions. However, models suffer from severe data scarcity and high annotation costs, severely…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Hongyu Zhu , Lin Chen , Xin Jin , Mingsheng Shang

Establishing correspondences between images remains a challenging task, especially under large appearance changes due to different viewpoints or intra-class variations. In this work, we introduce a strong semantic image matching learner,…

Computer Vision and Pattern Recognition · Computer Science 2022-05-25 Seungwook Kim , Juhong Min , Minsu Cho

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

Multimodal semantic communication has great potential to enhance downstream task performance by integrating complementary information across modalities. This paper introduces ProMSC-MIS, a novel Prompt-based Multimodal Semantic…

Multimedia · Computer Science 2025-08-28 Haoshuo Zhang , Yufei Bo , Meixia Tao

Existing text-driven infrared and visible image fusion approaches often rely on textual information at the sentence level, which can lead to semantic noise from redundant text and fail to fully exploit the deeper semantic value of textual…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Wenyu Shao , Hongbo Liu , Yunchuan Ma , Ruili Wang

Labeling semantic segmentation datasets is a costly and laborious process if compared with tasks like image classification and object detection. This is especially true for remote sensing applications that not only work with extremely high…

Computer Vision and Pattern Recognition · Computer Science 2021-08-27 Matheus Barros Pereira , Jefersson Alex dos Santos

Multispectral imaging (MSI) plays a critical role in material classification, environmental monitoring, and remote sensing. However, MSI sensors typically have wavelength-dependent resolution, which limits downstream analysis. MSI…

Image and Video Processing · Electrical Eng. & Systems 2026-03-10 Haley Duba-Sullivan , Emma J. Reid , Sophie Voisin , Charles A. Bouman , Gregery T. Buzzard

Synthetic Aperture Radar (SAR) and optical imagery provide complementary strengths that constitute the critical foundation for transcending single-modality constraints and facilitating cross-modal collaborative processing and intelligent…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Peihao Wu , Yongxiang Yao , Yi Wan , Wenfei Zhang , Ruipeng Zhao , Jiayuan Li , Yongjun Zhang

Remote Sensing Image Super-Resolution (RSISR) reconstructs high-resolution (HR) remote sensing images from low-resolution inputs to support fine-grained ground object interpretation. Existing methods face three key challenges: (1)…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Yide Liu , Haijiang Sun , Xiaowen Zhang , Qiaoyuan Liu , Zhouchang Chen , Chongzhuo Xiao

Cross-modality fusing complementary information of multispectral remote sensing image pairs can improve the perception ability of detection algorithms, making them more robust and reliable for a wider range of applications, such as…

Computer Vision and Pattern Recognition · Computer Science 2021-12-07 Qingyun Fang , Zhaokui Wang

Multimodal machine translation (MMT) aims to improve neural machine translation (NMT) with additional visual information, but most existing MMT methods require paired input of source sentence and image, which makes them suffer from shortage…

Computation and Language · Computer Science 2022-03-22 Qingkai Fang , Yang Feng

Visual grounding aims to localize the object referred to in an image based on a natural language query. Although progress has been made recently, accurately localizing target objects within multiple-instance distractions (multiple objects…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Minghang Zheng , Jiahua Zhang , Qingchao Chen , Yuxin Peng , Yang Liu

The abundance of multimodal data (e.g. social media posts) has inspired interest in cross-modal retrieval methods. Popular approaches rely on a variety of metric learning losses, which prescribe what the proximity of image and text should…

Computer Vision and Pattern Recognition · Computer Science 2020-09-24 Christopher Thomas , Adriana Kovashka

In this work, we investigate various methods to deal with semantic labeling of very high resolution multi-modal remote sensing data. Especially, we study how deep fully convolutional networks can be adapted to deal with multi-modal and…

Neural and Evolutionary Computing · Computer Science 2017-11-27 Nicolas Audebert , Bertrand Le Saux , Sébastien Lefèvre

The training of large multimodal models fundamentally relies on massive image-text datasets, which inevitably incur prohibitive computational overhead. Dataset selection offers a promising paradigm by identifying a highly informative…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Boran Zhao , Hetian Liu , Zhenxian Hu , Yuqing Yuan , Yu Yan , Pengju Ren

Synthetic Aperture Radar (SAR) and optical image registration is essential for remote sensing data fusion, with applications in military reconnaissance, environmental monitoring, and disaster management. However, challenges arise from…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Wenfei Zhang , Ruipeng Zhao , Yongxiang Yao , Yi Wan , Peihao Wu , Jiayuan Li , Yansheng Li , Yongjun Zhang

Recent advancements in multimodal large language models (MLLMs) have made significant progress in integrating information across various modalities, yet real-world applications in educational and scientific domains remain challenging. This…

Computation and Language · Computer Science 2024-11-15 Minghan Wang , Yuxia Wang , Thuy-Trang Vu , Ehsan Shareghi , Gholamreza Haffari

Remote Sensing Visual Question Answering (RSVQA) is a task that extracts information from satellite images to answer questions in natural language, aiding image interpretation. While several methods exist for optical images with varying…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Lucrezia Tosato , Flora Weissgerber , Laurent Wendling , Sylvain Lobry

This manuscript explores multimodal alignment, translation, fusion, and transference to enhance machine understanding of complex inputs. We organize the work into five chapters, each addressing unique challenges in multimodal machine…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Gorjan Radevski

Multi-modal data fusion has recently been shown promise in classification tasks in remote sensing. Optical data and radar data, two important yet intrinsically different data sources, are attracting more and more attention for potential…

Computer Vision and Pattern Recognition · Computer Science 2020-01-08 Jingliang Hu , Danfeng Hong , Xiao Xiang Zhu
‹ Prev 1 8 9 10 Next ›