中文
相关论文

相关论文: TAR: Text Semantic Assisted Cross-modal Image Regi…

200 篇论文

This paper presents a multimodal framework that attempts to unify visual understanding and generation within a shared discrete semantic representation. At its core is the Text-Aligned Tokenizer (TA-Tok), which converts images into discrete…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Jiaming Han , Hao Chen , Yang Zhao , Hanyu Wang , Qi Zhao , Ziyan Yang , Hao He , Xiangyu Yue , Lu Jiang

Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability. Although recent advances in self-supervised learning and masked image modeling (MIM) have enabled SAR foundation models, these…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Qiwei Ma , Xukun Lu , Wang Liu , Puhong Duan , Xudong Kang , Shutao Li

Three-dimensional synthetic aperture radar (3D SAR) is an advanced active microwave imaging technology widely utilized in remote sensing area. To achieve high-resolution 3D imaging,3D SAR requires observations from multiple aspects and…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Da Li , Guoqiang Zhao , Chen Yao , Kaiqiang Zhu , Houjun Sun , Jiacheng Bao , Maokun Li

Though significant progress in human pose and shape recovery from monocular RGB images has been made in recent years, obtaining 3D human motion with high accuracy and temporal consistency from videos remains challenging. Existing…

计算机视觉与模式识别 · 计算机科学 2023-11-17 Ming Chen , Yan Zhou , Weihua Jian , Pengfei Wan , Zhongyuan Wang

Registration of optical and synthetic aperture radar (SAR) remote sensing images serves as a critical foundation for image fusion and visual navigation tasks. This task is particularly challenging because of their modal discrepancy,…

图像与视频处理 · 电气工程与系统科学 2025-11-04 Zixuan Sun , Shuaifeng Zhi , Ruize Li , Jingyuan Xia , Yongxiang Liu , Weidong Jiang

The recently proposed DEtection TRansformer (DETR) has established a fully end-to-end paradigm for object detection. However, DETR suffers from slow training convergence, which hinders its applicability to various detection tasks. We…

计算机视觉与模式识别 · 计算机科学 2023-02-07 Gongjie Zhang , Zhipeng Luo , Jiaxing Huang , Shijian Lu , Eric P. Xing

The challenge of Multimodal Deformable Image Registration (MDIR) lies in the conversion and alignment of features between images of different modalities. Generative models (GMs) cannot retain the necessary information enough from the source…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Mingrui Ma , Weijie Wang , Jie Ning , Jianfeng He , Nicu Sebe , Bruno Lepri

Research on the intelligent interpretation of all-weather, all-time Synthetic Aperture Radar (SAR) is crucial for advancing remote sensing applications. In recent years, although Visual Language Models (VLMs) have demonstrated strong…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Xiaokun Zhang , Yi Yang , Ziqi Ye , Baiyun , Xiaorong Guo , Qingchen Fang , Ruyi Zhang , Xinpeng Zhou , Haipeng Wang

While image registration has been studied in remote sensing community for decades, registering multimodal data [e.g., optical, LiDAR, SAR, and map] remains a challenging problem because of significant nonlinear intensity differences between…

计算机视觉与模式识别 · 计算机科学 2021-04-01 Yuanxin Ye , Lorenzo Bruzzone , Jie Shan , Francesca Bovolo , Qing Zhu

Human action understanding is crucial for the advancement of multimodal systems. While recent developments, driven by powerful large language models (LLMs), aim to be general enough to cover a wide range of categories, they often overlook…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Yongle Huang , Haodong Chen , Zhenbang Xu , Zihan Jia , Haozhou Sun , Dian Shao

Multi-modal image registration is a challenging problem that is also an important clinical task for many real applications and scenarios. As a first step in analysis, deformable registration among different image modalities is often…

图像与视频处理 · 电气工程与系统科学 2020-07-21 Fengze Liu , Jinzheng Cai , Yuankai Huo , Chi-Tung Cheng , Ashwin Raju , Dakai Jin , Jing Xiao , Alan Yuille , Le Lu , ChienHung Liao , Adam P Harrison

Synthetic aperture radar (SAR) imaging technology is commonly used to provide 24-hour all-weather earth observation. However, it still has some drawbacks in SAR target classification, especially in fine-grained classification of aircraft:…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Bingying Yue , Jianhao Li , Hao Shi , Yupei Wang , Honghu Zhong

Recent work on action recognition leverages 3D features and textual information to achieve state-of-the-art performance. However, most of the current few-shot action recognition methods still rely on 2D frame-level representations, often…

计算机视觉与模式识别 · 计算机科学 2023-11-13 Yutao Tang , Benjamin Bejar , Rene Vidal

The effective combination of the complementary information provided by the huge amount of unlabeled multi-sensor data (e.g., Synthetic Aperture Radar (SAR) and optical images) is a critical topic in remote sensing. Recently, contrastive…

图像与视频处理 · 电气工程与系统科学 2021-10-11 Yuxing Chen , Lorenzo Bruzzone

Image-text retrieval (ITR) is a challenging task in the field of multimodal information processing due to the semantic gap between different modalities. In recent years, researchers have made great progress in exploring the accurate…

计算机视觉与模式识别 · 计算机科学 2022-12-19 Jie Guo , Meiting Wang , Yan Zhou , Bin Song , Yuhao Chi , Wei Fan , Jianglong Chang

Optical and Synthetic Aperture Radar (SAR) fusion-based object detection has attracted significant research interest in remote sensing, as these modalities provide complementary information for all-weather monitoring. However, practical…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Zhicheng Zhao , Yuancheng Xu , Andong Lu , Chenglong Li , Jin Tang

For many computer vision applications such as image captioning, visual question answering, and person search, learning discriminative feature representations at both image and text level is an essential yet challenging problem. Its…

计算机视觉与模式识别 · 计算机科学 2019-08-29 Nikolaos Sarafianos , Xiang Xu , Ioannis A. Kakadiaris

Automatic radiology report generation has attracted enormous research interest due to its practical value in reducing the workload of radiologists. However, simultaneously establishing global correspondences between the image (e.g., Chest…

计算机视觉与模式识别 · 计算机科学 2023-07-19 Yaowei Li , Bang Yang , Xuxin Cheng , Zhihong Zhu , Hongxiang Li , Yuexian Zou

Previous multimodal sentence representation learning methods have achieved impressive performance. However, most approaches focus on aligning images and text at a coarse level, facing two critical challenges:cross-modal misalignment bias…

计算与语言 · 计算机科学 2025-07-02 Kang He , Yuzhe Ding , Haining Wang , Fei Li , Chong Teng , Donghong Ji

Image registration under domain shift remains a fundamental challenge in computer vision and medical imaging: when source and target images exhibit systematic intensity differences, the brightness constancy assumption underlying…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Jiahao Qin , Yiwen Wang