English
Related papers

Related papers: CrossModalityDiffusion: Multi-Modal Novel View Syn…

200 papers

Recently, large-scale diffusion models, e.g., Stable diffusion and DallE2, have shown remarkable results on image synthesis. On the other hand, large-scale cross-modal pre-trained models (e.g., CLIP, ALIGN, and FILIP) are competent for…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Runhui Huang , Jianhua Han , Guansong Lu , Xiaodan Liang , Yihan Zeng , Wei Zhang , Hang Xu

Multi-modality image fusion aims at fusing modality-specific (complementarity) and modality-shared (correlation) information from multiple source images. To tackle the problem of the neglect of inter-feature relationships, high-frequency…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Xiaoli Zhang , Liying Wang , Libo Zhao , Xiongfei Li , Siwei Ma

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a shared diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Junyi Chen , Tong He , Zhoujie Fu , Pengfei Wan , Kun Gai , Weicai Ye

In this paper, we focus on 3D scene inpainting, where parts of an input image set, captured from different viewpoints, are masked out. The main challenge lies in generating plausible image completions that are geometrically consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Ahmad Salimi , Tristan Aumentado-Armstrong , Marcus A. Brubaker , Konstantinos G. Derpanis

Generating consistent multi-view images from a single image remains challenging. Lack of spatial consistency often degrades 3D mesh quality in surface reconstruction. To address this, we propose LoomNet, a novel multi-view diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Giulio Federico , Fabio Carrara , Claudio Gennaro , Giuseppe Amato , Marco Di Benedetto

Despite advances in object detection, aerial imagery remains a challenging domain, as models often fail to generalize across variations in spatial resolution, scene composition, and semantic label coverage. Differences in geographic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Pourya Shamsolmoali , Masoumeh Zareapoor , Michael Felsberg , Nick Pears , Yue Lu

The combination of LiDAR and camera modalities is proven to be necessary and typical for 3D object detection according to recent studies. Existing fusion strategies tend to overly rely on the LiDAR modal in essence, which exploits the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Yang Yang , Weijie Ma , Hao Chen , Linlin Ou , Xinyi Yu

The recent popularity of text-to-image diffusion models (DM) can largely be attributed to the intuitive interface they provide to users. The intended generation can be expressed in natural language, with the model producing faithful…

Object detection in Remote Sensing Images (RSI) is a critical task for numerous applications in Earth Observation (EO). Differing from object detection in natural images, object detection in remote sensing images faces challenges of…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Bissmella Bahaduri , Zuheng Ming , Fangchen Feng , Anissa Mokraou

Cross-modality data translation has attracted great interest in image computing. Deep generative models (\textit{e.g.}, GANs) show performance improvement in tackling those problems. Nevertheless, as a fundamental challenge in image…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Zihao Wang , Yingyu Yang , Maxime Sermesant , Hervé Delingette , Ona Wu

Multi-modal image fusion (MMIF) maps useful information from various modalities into the same representation space, thereby producing an informative fused image. However, the existing fusion algorithms tend to symmetrically fuse the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Jingxue Huang , Xilai Li , Tianshu Tan , Xiaosong Li , Tao Ye

Multi-modal image fusion aims to integrate complementary information from multiple source images to produce high-quality fused images with enriched content. Although existing approaches based on state space model have achieved satisfied…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Yiming Sun , Zifan Ye , Qinghua Hu , Pengfei Zhu

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Zhaoyang Li , Dongjun Qian , Kai Su , Qishuai Diao , Xiangyang Xia , Chang Liu , Wenfei Yang , Tianzhu Zhang , Zehuan Yuan

In this work, we introduce Wonder3D, a novel method for efficiently generating high-fidelity textured meshes from single-view images.Recent methods based on Score Distillation Sampling (SDS) have shown the potential to recover 3D geometry…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Xiaoxiao Long , Yuan-Chen Guo , Cheng Lin , Yuan Liu , Zhiyang Dou , Lingjie Liu , Yuexin Ma , Song-Hai Zhang , Marc Habermann , Christian Theobalt , Wenping Wang

Denosing diffusion model, as a generative model, has received a lot of attention in the field of image generation recently, thanks to its powerful generation capability. However, diffusion models have not yet received sufficient research in…

Computer Vision and Pattern Recognition · Computer Science 2023-04-12 ZiHan Cao , ShiQi Cao , Xiao Wu , JunMing Hou , Ran Ran , Liang-Jian Deng

For better explore the relations of inter-modal and inner-modal, even in deep learning fusion framework, the concept of decomposition plays a crucial role. However, the previous decomposition strategies (base \& detail or low-frequency \&…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Hui Li , Haolong Ma , Chunyang Cheng , Zhongwei Shen , Xiaoning Song , Xiao-Jun Wu

Humans perceive the world through multimodal cues to understand and interact with the environment. Learning a scene representation for multiple modalities enhances comprehension of the physical world. However, modality conflicts, arising…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Zhifeng Gu , Bing Wang

Color plays an important role in human visual perception, reflecting the spectrum of objects. However, the existing infrared and visible image fusion methods rarely explore how to handle multi-spectral/channel data directly and achieve high…

Computer Vision and Pattern Recognition · Computer Science 2024-10-28 Jun Yue , Leyuan Fang , Shaobo Xia , Yue Deng , Jiayi Ma

The visual entities in cross-view images exhibit drastic domain changes due to the difference in viewpoints each set of images is captured from. Existing state-of-the-art methods address the problem by learning view-invariant descriptors…

Computer Vision and Pattern Recognition · Computer Science 2019-08-12 Krishna Regmi , Mubarak Shah

Cross-modality medical image synthesis is a critical topic and has the potential to facilitate numerous applications in the medical imaging field. Despite recent successes in deep-learning-based generative models, most current medical image…

Image and Video Processing · Electrical Eng. & Systems 2023-07-20 Lingting Zhu , Zeyue Xue , Zhenchao Jin , Xian Liu , Jingzhen He , Ziwei Liu , Lequan Yu
‹ Prev 1 3 4 5 6 7 10 Next ›