English
Related papers

Related papers: EarthBridge: A Solution for 4th Multi-modal Aerial…

200 papers

The purpose of this study is to apply and evaluate out-of-the-box deep learning frameworks for the crossMoDA challenge. We use the CUT model, a model for unpaired image-to-image translation based on patchwise contrastive learning and…

Image and Video Processing · Electrical Eng. & Systems 2021-12-09 Jae Won Choi

In clinical practice, well-aligned multi-modal images, such as Magnetic Resonance (MR) and Computed Tomography (CT), together can provide complementary information for image-guided therapies. Multi-modal image registration is essential for…

Computer Vision and Pattern Recognition · Computer Science 2022-04-29 Zekang Chen , Jia Wei , Rui Li

The volume of unlabelled Earth observation (EO) data is huge, but many important applications lack labelled training data. However, EO data offers the unique opportunity to pair data from different modalities and sensors automatically based…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Vishal Nedungadi , Ankit Kariryaa , Stefan Oehmcke , Serge Belongie , Christian Igel , Nico Lang

We study the problem of multimodal fusion in this paper. Recent exchanging-based methods have been proposed for vision-vision fusion, which aim to exchange embeddings learned from one modality to the other. However, most of them project…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Renyu Zhu , Chengcheng Han , Yong Qian , Qiushi Sun , Xiang Li , Ming Gao , Xuezhi Cao , Yunsen Xian

Unsupervised image-to-image translation aims at learning a joint distribution of images in different domains by using images from the marginal distributions in individual domains. Since there exists an infinite set of joint distributions…

Computer Vision and Pattern Recognition · Computer Science 2018-07-24 Ming-Yu Liu , Thomas Breuel , Jan Kautz

Diffusion bridge models have recently become a powerful tool in the field of generative modeling. In this work, we leverage their power to address another important problem in machine learning and information theory, the estimation of the…

Machine Learning · Computer Science 2026-03-02 Sergei Kholkin , Ivan Butakov , Evgeny Burnaev , Nikita Gushchin , Alexander Korotin

In this paper, we introduce Latent Bridge Matching (LBM), a new, versatile and scalable method that relies on Bridge Matching in a latent space to achieve fast image-to-image translation. We show that the method can reach state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Clément Chadebec , Onur Tasar , Sanjeev Sreetharan , Benjamin Aubin

Unmanned aerial vehicle (UAV) detection and aerial object recognition are critical for modern surveillance and security, prompting a need for robust systems that overcome limitations of single-modality approaches. This research addresses…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Mauro Larrat , Claudomiro Sales

Text-to-image generation has made significant advancements with the introduction of text-to-image diffusion models. These models typically consist of a language model that interprets user prompts and a vision model that generates…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Shihao Zhao , Shaozhe Hao , Bojia Zi , Huaizhe Xu , Kwan-Yee K. Wong

Semantic communication has undergone considerable evolution due to the recent rapid development of artificial intelligence (AI), significantly enhancing both communication robustness and efficiency. Despite these advancements, most current…

Image and Video Processing · Electrical Eng. & Systems 2024-05-24 Jiarun Ding , Peiwen Jiang , Chao-Kai Wen , Shi Jin

The Dual Diffusion Implicit Bridge (DDIB) is an emerging image-to-image (I2I) translation method that preserves cycle consistency while achieving strong flexibility. It links two independently trained diffusion models (DMs) in the source…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Zhanpeng Wang , Shuting Cao , Yuhang Lu , Yuhan Li , Na Lei , Zhongxuan Luo

Aerial-ground localization is difficult due to large viewpoint and modality gaps between ground-level LiDAR and overhead imagery. We propose TransLocNet, a cross-modal attention framework that fuses LiDAR geometry with aerial semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Phu Pham , Damon Conover , Aniket Bera

In this work, we explore neat yet effective Transformer-based frameworks for visual grounding. The previous methods generally address the core problem of visual grounding, i.e., multi-modal fusion and reasoning, with manually-designed…

Computer Vision and Pattern Recognition · Computer Science 2022-06-15 Jiajun Deng , Zhengyuan Yang , Daqing Liu , Tianlang Chen , Wengang Zhou , Yanyong Zhang , Houqiang Li , Wanli Ouyang

In this paper, we propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, SVBRDF, and 3D spatially-varying lighting. While multi-view images have been widely used for object-level…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 JunYong Choi , SeokYeong Lee , Haesol Park , Seung-Won Jung , Ig-Jae Kim , Junghyun Cho

In this work, we develop methods for few-shot image classification from a new perspective of optimal matching between image regions. We employ the Earth Mover's Distance (EMD) as a metric to compute a structural distance between dense image…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Chi Zhang , Yujun Cai , Guosheng Lin , Chunhua Shen

Image-to-image translation aims to learn the mapping between two visual domains. There are two main challenges for many applications: 1) the lack of aligned training pairs and 2) multiple possible outputs from a single input image. In this…

Computer Vision and Pattern Recognition · Computer Science 2018-08-03 Hsin-Ying Lee , Hung-Yu Tseng , Jia-Bin Huang , Maneesh Kumar Singh , Ming-Hsuan Yang

Detecting visual relationships, i.e. <Subject, Predicate, Object> triplets, is a challenging Scene Understanding task approached in the past via linguistic priors or spatial information in a single feature branch. We introduce a new deeply…

Computer Vision and Pattern Recognition · Computer Science 2019-02-18 Nikolaos Gkanatsios , Vassilis Pitsikalis , Petros Koutras , Athanasia Zlatintsi , Petros Maragos

One task that is often discussed in a computer vision is the mapping of an image from one domain to a corresponding image in another domain known as image-to-image translation. Currently there are several approaches solving this task. In…

Machine Learning · Computer Science 2020-03-23 Magda Friedjungová , Daniel Vašata , Tomáš Chobola , Marcel Jiřina

We introduce Vision Bridge Transformer (ViBT), a large-scale instantiation of Brownian Bridge Models designed for conditional generation. Unlike traditional diffusion models that transform noise into data, Bridge Models directly model the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zhenxiong Tan , Zeqing Wang , Xingyi Yang , Songhua Liu , Xinchao Wang

This position paper argues that the next generation of artificial intelligence in meteorological and climate sciences must transition from fragmented hybrid heuristics toward a unified paradigm of physics-guided multimodal transformers.…

Machine Learning · Computer Science 2026-01-29 Jing Han , Hanting Chen , Kai Han , Xiaomeng Huang , Wenjun Xu , Dacheng Tao , Ping Zhang
‹ Prev 1 4 5 6 7 8 10 Next ›