English
Related papers

Related papers: CM-Diff: A Single Generative Network for Bidirecti…

200 papers

Multimodal intent understanding is a significant research area that requires effective leveraging of multiple modalities to analyze human language. Existing methods face two main challenges in this domain. Firstly, they have limitations in…

Multimedia · Computer Science 2025-05-26 Hanlei Zhang , Qianrui Zhou , Hua Xu , Jianhua Su , Roberto Evans , Kai Gao

Multimodal medical images play a crucial role in the precise and comprehensive clinical diagnosis. Diffusion model is a powerful strategy to synthesize the required medical images. However, existing approaches still suffer from the problem…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Jiahua Xu , Dawei Zhou , Lei Hu , Zaiyi Liu , Nannan Wang , Xinbo Gao

Multimodal medical image fusion (MMIF) extracts the most meaningful information from multiple source images, enabling a more comprehensive and accurate diagnosis. Achieving high-quality fusion results requires a careful balance of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Dan He , Weisheng Li , Guofen Wang , Yuping Huang , Shiqiang Liu

The capability to jointly process multi-modal information is becoming an essential task. However, the limited number of paired multi-modal data and the large computational requirements in multi-modal learning hinder the development. We…

Computation and Language · Computer Science 2025-06-09 Minsu Kim , Jee-weon Jung , Hyeongseop Rha , Soumi Maiti , Siddhant Arora , Xuankai Chang , Shinji Watanabe , Yong Man Ro

Visible thermal person re-identification (VT-ReID) suffers from the inter-modality discrepancy and intra-identity variations. Distribution alignment is a popular solution for VT-ReID, which, however, is usually restricted to the influence…

Computer Vision and Pattern Recognition · Computer Science 2022-03-04 Yongguo Ling , Zhun Zhong , Donglin Cao , Zhiming Luo , Yaojin Lin , Shaozi Li , Nicu Sebe

This study introduces an efficient and effective method, MeDM, that utilizes pre-trained image Diffusion Models for video-to-video translation with consistent temporal flow. The proposed framework can render videos from scene position…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Ernie Chu , Tzuhsuan Huang , Shuo-Yen Lin , Jun-Cheng Chen

For better explore the relations of inter-modal and inner-modal, even in deep learning fusion framework, the concept of decomposition plays a crucial role. However, the previous decomposition strategies (base \& detail or low-frequency \&…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Hui Li , Haolong Ma , Chunyang Cheng , Zhongwei Shen , Xiaoning Song , Xiao-Jun Wu

Language-guided image generation has achieved great success nowadays by using diffusion models. However, texts can be less detailed to describe highly-specific subjects such as a particular dog or a certain car, which makes pure…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Yiyang Ma , Huan Yang , Wenjing Wang , Jianlong Fu , Jiaying Liu

Processing and fusing information among multi-modal is a very useful technique for achieving high performance in many computer vision problems. In order to tackle multi-modal information more effectively, we introduce a novel framework for…

Computer Vision and Pattern Recognition · Computer Science 2019-05-01 Dong Wang , Yuan Yuan , Qi Wang

Unsupervised image-to-image translation tasks aim to find a mapping between a source domain X and a target domain Y from unpaired training data. Contrastive learning for Unpaired image-to-image Translation (CUT) yields state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Junlin Han , Mehrdad Shoeiby , Lars Petersson , Mohammad Ali Armin

Image compression technology eliminates redundant information to enable efficient transmission and storage of images, serving both machine vision and human visual perception. For years, image coding focused on human perception has been…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Takahiro Shindo , Yui Tatsumi , Taiju Watanabe , Hiroshi Watanabe

Cross-modal retrieval has drawn wide interest for retrieval across different modalities of data. However, existing methods based on DNN face the challenge of insufficient cross-modal training data, which limits the training effectiveness…

Multimedia · Computer Science 2017-08-16 Xin Huang , Yuxin Peng , Mingkuan Yuan

This paper addresses the problem of semi-supervised transfer learning with limited cross-modality data in remote sensing. A large amount of multi-modal earth observation images, such as multispectral imagery (MSI) or synthetic aperture…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Danfeng Hong , Naoto Yokoya , Gui-Song Xia , Jocelyn Chanussot , Xiao Xiang Zhu

We introduce a multi-modal diffusion model tailored for the bi-directional conditional generation of video and audio. We propose a joint contrastive training loss to improve the synchronization between visual and auditory occurrences. We…

Machine Learning · Computer Science 2024-10-10 Ruihan Yang , Hannes Gamper , Sebastian Braun

Diffusion Policy (DP) has attracted significant attention as an effective method for policy representation due to its capacity to model multi-distribution dynamics. However, current DPs are often based on a single visual modality (e.g., RGB…

Robotics · Computer Science 2025-03-18 Jiahang Cao , Qiang Zhang , Hanzhong Guo , Jiaxu Wang , Hao Cheng , Renjing Xu

Multimodal Machine Translation (MMT) has demonstrated the significant help of visual information in machine translation. However, existing MMT methods face challenges in leveraging the modality gap by enforcing rigid visual-linguistic…

Computation and Language · Computer Science 2025-10-09 Jiafeng Xiong , Yuting Zhao

Modality translation is inherently under-constrained, as multiple cross-modal mappings may yield the same marginals. Recent work has shown that diffusion bridges are effective for this task. However, most existing approaches rely on fully…

Machine Learning · Computer Science 2026-05-13 Eitan Kosman , Gabriele Serussi , Chaim Baskin

To tackle the challenge of vehicle re-identification (Re-ID) in complex lighting environments and diverse scenes, multi-spectral sources like visible and infrared information are taken into consideration due to their excellent complementary…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Aihua Zheng , Xianpeng Zhu , Zhiqi Ma , Chenglong Li , Jin Tang , Jixin Ma

Multi-modal learning is typically performed with network architectures containing modality-specific layers and shared layers, utilizing co-registered images of different modalities. We propose a novel learning scheme for unpaired…

Computer Vision and Pattern Recognition · Computer Science 2020-01-10 Qi Dou , Quande Liu , Pheng Ann Heng , Ben Glocker

Multi-modal brain images from MRI scans are widely used in clinical diagnosis to provide complementary information from different modalities. However, obtaining fully paired multi-modal images in practice is challenging due to various…

Image and Video Processing · Electrical Eng. & Systems 2024-04-25 Chuan Huang , Jia Wei , Rui Li