English
Related papers

Related papers: Inconsistency-aware Multimodal Schr\"odinger Bridg…

200 papers

Schrodinger Bridges (SBs) are diffusion processes that steer, in finite time, a given initial distribution to another final one while minimizing a suitable cost functional. Although various methods for computing SBs have recently been…

Machine Learning · Computer Science 2025-10-15 George Rapakoulias , Ali Reza Pedram , Fengjiao Liu , Lingjiong Zhu , Panagiotis Tsiotras

Recent advancements in unpaired dehazing, particularly those using GANs, show promising performance in processing real-world hazy images. However, these methods tend to face limitations due to the generator's limited transport mapping…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Yunwei Lan , Zhigao Cui , Xin Luo , Chang Liu , Nian Wang , Menglin Zhang , Yanzhao Su , Dong Liu

Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust perception. However, notable limitations remain: (1) existing methods often use text only as…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Jiaqi Wu , Zhen Wang , Enhao Huang , Kangqing Shen , Yulin Wang , Yang Yue , Yifan Pu , Gao Huang

Change detection plays a fundamental role in Earth observation for analyzing temporal iterations over time. However, recent studies have largely neglected the utilization of multimodal data that presents significant practical and technical…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Biyuan Liu , Huaixin Chen , Kun Li , Michael Ying Yang

In this paper, we investigate the multi-marginal Schrodinger bridge (MSB) problem whose marginal constraints are marginal distributions of a stochastic differential equation (SDE) with a constant diffusion coefficient, and with time…

Probability · Mathematics 2025-07-15 Rentian Yao , Young--Heon Kim , Geoffrey Schiebinger

Audio-visual temporal deepfake localization under the content-driven partial manipulation remains a highly challenging task. In this scenario, the deepfake regions are usually only spanning a few frames, with the majority of the rest…

Unsupervised domain adaptation (UDA) methods effectively bridge domain gaps but become struggled when the source and target domains belong to entirely distinct modalities. To address this limitation, we propose a novel setting called…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Jiawen Yang , Shuhao Chen , Yucong Duan , Ke Tang , Yu Zhang

Indirect structural health monitoring (iSHM) for broken rail detection using onboard sensors presents a cost-effective paradigm for railway track assessment, yet reliably detecting small, transient anomalies (2-10 cm) remains a significant…

Machine Learning · Computer Science 2025-10-10 Sizhe Ma , Katherine A. Flanigan , Mario Bergés , James D. Brooks

Multimodal deep learning harnesses diverse imaging modalities, such as MRI sequences, to enhance diagnostic accuracy in medical imaging. A key challenge is determining the optimal timing for integrating these modalities-specifically,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Valerio Guarrasi , Klara Mogensen , Sara Tassinari , Sara Qvarlander , Paolo Soda

Multimodal semantic segmentation has shown great potential in leveraging complementary information across diverse sensing modalities. However, existing approaches often rely on carefully designed fusion strategies that either use…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Zelin Zhang , Kedi Li , Huiqi Liang , Tao Zhang , Chuanzhi Xu

Recent advances in generative modeling have positioned diffusion models as state-of-the-art tools for sampling from complex data distributions. While these models have shown remarkable success across single-modality domains such as images…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Nimrod Berman , Omkar Joglekar , Eitan Kosman , Dotan Di Castro , Omri Azencot

The rapid advancement of generative adversarial networks (GANs) and diffusion models has enabled the creation of highly realistic deepfake content, posing significant threats to digital trust across audio-visual domains. While unimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Chende Zheng , Ruiqi Suo , Zhoulin Ji , Jingyi Deng , Fangbin Yi , Chenhao Lin , Chao Shen

In this paper, we present novel synthetic training data called self-blended images (SBIs) to detect deepfakes. SBIs are generated by blending pseudo source and target images from single pristine images, reproducing common forgery artifacts…

Computer Vision and Pattern Recognition · Computer Science 2022-04-19 Kaede Shiohara , Toshihiko Yamasaki

Multi-modal stance detection (MSD) aims to determine an author's stance toward a given target using both textual and visual content. While recent methods leverage multi-modal fusion and prompt-based learning, most fail to distinguish…

Multimedia · Computer Science 2026-01-30 Zhiyu Xie , Fuqiang Niu , Genan Dai , Qianlong Wang , Li Dong , Bowen Zhang , Hu Huang

Audio-visual deepfake detection scrutinizes manipulations in public video using complementary multimodal cues. Current methods, which train on fused multimodal data for multimodal targets face challenges due to uncertainties and…

Multimedia · Computer Science 2024-01-12 Heqing Zou , Meng Shen , Yuchen Hu , Chen Chen , Eng Siong Chng , Deepu Rajan

Diffusion models have become a leading paradigm for image super-resolution (SR), but existing methods struggle to guarantee both the high-frequency perceptual quality and the low-frequency structural fidelity of generated images. Although…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Hexin Zhang , Dong Li , Jie Huang , Bingzhou Wang , Xueyang Fu , Zhengjun Zha

Image Splicing Localization (ISL) is a fundamental yet challenging task in digital forensics. Although current approaches have achieved promising performance, the edge information is insufficiently exploited, resulting in poor integrality…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Yakun Niu , Pei Chen , Lei Zhang , Hongjian Yin , Qi Chang

Remote sensing change detection is vital for monitoring environmental and urban transformations but faces challenges like manual feature extraction and sensitivity to noise. Traditional methods and early deep learning models, such as…

Generating samples from complex and high-dimensional distributions is ubiquitous in various scientific fields of statistical physics, Bayesian inference, scientific computing and machine learning. Very recently, Huang et al. (IEEE Trans.…

Numerical Analysis · Mathematics 2026-01-01 Xiaojie Wang , Xiaoyan Zhang

In recent advancements in proton therapy, MR-based treatment planning is gaining momentum to minimize additional radiation exposure compared to traditional CT-based methods. This transition highlights the critical need for accurate MR-to-CT…

Medical Physics · Physics 2024-07-02 Muheng Li , Xia Li , Sairos Safai , Damien Weber , Antony Lomax , Ye Zhang