English
Related papers

Related papers: XoFTR: Cross-modal Feature Matching Transformer

200 papers

Multimode fiber (MMF) imaging aided by machine learning holds promise for numerous applications, including medical endoscopy. A key challenge for this technology is the sensitivity of modal transmission characteristics to environmental…

Action recognition from multi-modal and multi-view observations holds significant potential for applications in surveillance, robotics, and smart environments. However, existing methods often fall short of addressing real-world challenges…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Trung Thanh Nguyen , Yasutomo Kawanishi , Vijay John , Takahiro Komamizu , Ichiro Ide

Vision-language retrieval is an important multi-modal learning topic, where the goal is to retrieve the most relevant visual candidate for a given text query. Recently, pre-trained models, e.g., CLIP, show great potential on retrieval…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Haojun Jiang , Jianke Zhang , Rui Huang , Chunjiang Ge , Zanlin Ni , Shiji Song , Gao Huang

Reliable perception is essential for autonomous driving systems to operate safely under diverse real-world traffic conditions. However, camera- and LiDAR-based perception systems suffer from performance degradation under adverse weather and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yue Sun , Yeqiang Qian , Zhe Wang , Tianhui Li , Chunxiang Wang , Ming Yang

Image feature matching, a foundational task in computer vision, remains challenging for multimodal image applications, often necessitating intricate training on specific datasets. In this paper, we introduce a Unified Feature Matching…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Yide Di , Yun Liao , Hao Zhou , Kaijun Zhu , Qing Duan , Junhui Liu , Mingyu Lu

We address the problem of visible-infrared person re-identification (VI-reID), that is, retrieving a set of person images, captured by visible or infrared cameras, in a cross-modal setting. Two main challenges in VI-reID are intra-class…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Hyunjong Park , Sanghoon Lee , Junghyup Lee , Bumsub Ham

Identity-invariant facial expression recognition (FER) has been one of the challenging computer vision tasks. Since conventional FER schemes do not explicitly address the inter-identity variation of facial expressions, their neural network…

Computer Vision and Pattern Recognition · Computer Science 2022-09-27 Daeha Kim , Byung Cheol Song

Robust feature representation plays significant role in visual tracking. However, it remains a challenging issue, since many factors may affect the experimental performance. The existing method which combine different features by setting…

Computer Vision and Pattern Recognition · Computer Science 2017-05-15 Yuqi Han , Chenwei Deng , Zengshuo Zhang , Jiatong Li , Baojun Zhao

Object detection in remote sensing imagery plays a vital role in various Earth observation applications. However, unlike object detection in natural scene images, this task is particularly challenging due to the abundance of small, often…

Computer Vision and Pattern Recognition · Computer Science 2024-09-16 Minh-Duc Vu , Zuheng Ming , Fangchen Feng , Bissmella Bahaduri , Anissa Mokraoui

Cross-modal object tracking is an important research topic in the field of information fusion, and it aims to address imaging limitations in challenging scenarios by integrating switchable visible and near-infrared modalities. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-25 Lei Liu , Chenglong Li , Futian Wang , Longfeng Shen , Jin Tang

Biological soft tissues encountered in clinical and pre-clinical imaging mainly consist of light element atoms, and their composition is nearly uniform with little density variation. Thus, x-ray attenuation imaging suffers from low image…

Medical Physics · Physics 2011-06-28 Wenxiang Cong , Atsushi Momose , Ge Wang

Viewport prediction is a crucial aspect of tile-based 360 video streaming system. However, existing trajectory based methods lack of robustness, also oversimplify the process of information construction and fusion between different modality…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Zhihao Zhang , Yiwei Chen , Weizhan Zhang , Caixia Yan , Qinghua Zheng , Qi Wang , Wangdu Chen

Multimodal visual information fusion aims to integrate the multi-sensor data into a single image which contains more complementary information and less redundant features. However the complementary information is hard to extract, especially…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Hui Li , Xiao-Jun Wu

With the rapid progression of deep learning technologies, multi-modality image fusion has become increasingly prevalent in object detection tasks. Despite its popularity, the inherent disparities in how different sources depict scene…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Xingyuan Li , Yang Zou , Jinyuan Liu , Zhiying Jiang , Long Ma , Xin Fan , Risheng Liu

This report introduces a solution to The task of RGB-TIR object detection from the perspective of unmanned aerial vehicles. Unlike traditional object detection methods, RGB-TIR object detection aims to utilize both RGB and TIR images for…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Xiangyu Wu , Jinling Xu , Longfei Huang , Yang Yang

Enhancing scene understanding in adverse visibility conditions remains a critical challenge for surveillance and autonomous navigation systems. Conventional imaging modalities, such as RGB and thermal infrared (MWIR / LWIR), when fused,…

Machine Learning · Computer Science 2025-11-25 Muhammad Ishfaq Hussain , Ma Van Linh , Zubia Naz , Unse Fatima , Yeongmin Ko , Moongu Jeon

Image registration is an essential process for aligning features of interest from multiple images. With the recent development of deep learning techniques, image registration approaches have advanced to a new level. In this work, we present…

Image and Video Processing · Electrical Eng. & Systems 2024-07-30 Ruixiong Wang , Alin Achim , Renata Raele-Rolfe , Qiao Tong , Dylan Bergen , Chrissy Hammond , Stephen Cross

To achieve accurate 3D object detection at a low cost for autonomous driving, many multi-camera methods have been proposed and solved the occlusion problem of monocular approaches. However, due to the lack of accurate estimated depth,…

Computer Vision and Pattern Recognition · Computer Science 2023-02-06 Ching-Yu Tseng , Yi-Rong Chen , Hsin-Ying Lee , Tsung-Han Wu , Wen-Chin Chen , Winston H. Hsu

Video transition effects are widely used in video editing to connect shots for creating cohesive and visually appealing videos. However, it is challenging for non-professionals to choose best transitions due to the lack of cinematographic…

Computer Vision and Pattern Recognition · Computer Science 2022-07-28 Yaojie Shen , Libo Zhang , Kai Xu , Xiaojie Jin

Infrared-visible image fusion aims to create an information-rich fused image by integrating the complementary thermal saliency from infrared sensing and fine textures from visible imaging. Such accurate fusion is essential for real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Zhenyu Sun , Luobin Zhang , Axi Niu , Haishen Wang , Qingsen Yan