中文
相关论文

相关论文: Inconsistency-aware Multimodal Schr\"odinger Bridg…

200 篇论文

Spatiotemporal image generation is a highly meaningful task, which can generate future scenes conditioned on given observations. However, existing change generation methods can only handle event-driven changes (e.g., new buildings) and fail…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Zhenghui Zhao , Chen Wu , Xiangyong Cao , Di Wang , Hongruixuan Chen , Datao Tang , Liangpei Zhang , Zhuo Zheng

Change detection (CD) in multitemporal remote sensing imagery presents significant challenges for fine-grained recognition, owing to heterogeneity and spatiotemporal misalignment. However, existing methodologies based on vision transformers…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Lei Ding , Tong Liu , Xuanguang Liu , Xiangyun Liu , Haitao Guo , Jun Lu

Video multimodal fusion aims to integrate multimodal signals in videos, such as visual, audio and text, to make a complementary prediction with multiple modalities contents. However, unlike other image-text multimodal tasks, video has…

计算与语言 · 计算机科学 2023-06-01 Shaoxiang Wu , Damai Dai , Ziwei Qin , Tianyu Liu , Binghuai Lin , Yunbo Cao , Zhifang Sui

Deepfake technology has rapidly advanced and poses significant threats to information integrity and trust in online multimedia. While significant progress has been made in detecting deepfakes, the simultaneous manipulation of audio and…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Christos Koutlis , Symeon Papadopoulos

Matrix pair beamformer (MPB) is a promising blind beamformer which exploits the temporal signature of the signal of interest (SOI) to acquire its spatial statistical information. It does not need any knowledge of directional information or…

计算工程、金融与科学 · 计算机科学 2010-10-04 Jian Wang , Jianshu Chen , Jian Yuan , Ning Ge , Shuangqing Wei

Subsurface defects such as delamination, voids, and honeycombing critically affect the durability of concrete bridge decks but are difficult to detect reliably using visual inspection or manual sounding. This paper presents a machine…

信号处理 · 电气工程与系统科学 2025-11-27 Yeswanth Ravichandran , Duoduo Liao , Charan Teja Kurakula

Standard multimodal self-supervised learning (SSL) algorithms regard cross-modal synchronization as implicit supervisory labels during pretraining, thus posing high requirements on the scale and quality of multimodal samples. These…

Diffusion bridges (DB) have emerged as a promising alternative to diffusion models for imaging inverse problems, achieving faster sampling by directly bridging low- and high-quality image distributions. While incorporating measurement…

图像与视频处理 · 电气工程与系统科学 2024-11-26 Yuyang Hu , Albert Peng , Weijie Gan , Ulugbek S. Kamilov

Recently, a series of papers proposed deep learning-based approaches to sample from target distributions using controlled diffusion processes, being trained only on the unnormalized target densities without access to samples. Building on…

机器学习 · 计算机科学 2024-05-24 Lorenz Richter , Julius Berner

Vehicles crossing bridge structures respond dynamically to the bridge's vibrations. An acceleration signal collected within a moving vehicle contains a trace of the bridge's structural response, but also includes other sources such as the…

信号处理 · 电气工程与系统科学 2020-03-05 Soheil Sadeghi Eshkevari , Thomas J. Matarazzo , Shamim N. Pakzad

Understanding indoor scenes is crucial for urban studies. Considering the dynamic nature of indoor environments, effective semantic segmentation requires both real-time operation and high accuracy.To address this, we propose AsymFormer, a…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Siqi Du , Weixi Wang , Renzhong Guo , Ruisheng Wang , Yibin Tian , Shengjun Tang

Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This…

音频与语音处理 · 电气工程与系统科学 2025-01-22 Julius Richter , Danilo de Oliveira , Timo Gerkmann

Diffusion models (DMs) have become the dominant paradigm of generative modeling in a variety of domains by learning stochastic processes from noise to data. Recently, diffusion denoising bridge models (DDBMs), a new formulation of…

机器学习 · 计算机科学 2024-11-01 Guande He , Kaiwen Zheng , Jianfei Chen , Fan Bao , Jun Zhu

Multimodal semantic segmentation has emerged as a powerful paradigm for enhancing scene understanding by leveraging complementary information from multiple sensing modalities (e.g., RGB, depth, and thermal). However, existing cross-modal…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Guoan Xu , Yang Xiao , Guangwei Gao , Dongchen Zhu , Guo-Jun Qi , Wenjing Jia

In the digital age, the emergence of deepfakes and synthetic media presents a significant threat to societal and political integrity. Deepfakes based on multi-modal manipulation, such as audio-visual, are more realistic and pose a greater…

声音 · 计算机科学 2024-08-08 Vinaya Sree Katamneni , Ajita Rattani

Accurate detection of road and bridge changes is crucial for urban planning and transportation management, yet presents unique challenges for general change detection (CD). Key difficulties arise from maintaining the continuity of roads and…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Qingling Shu , Sibao Chen , Xiao Wang , Zhihui You , Wei Lu , Jin Tang , Bin Luo

We investigate the unsupervised node classification problem on random hypergraphs under the non-uniform Hypergraph Stochastic Block Model (HSBM) with two equal-sized communities. In this model, edges appear independently with probabilities…

统计理论 · 数学 2025-12-01 Hai-Xiao Wang

Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanced methods such as BAGEL and LMFusion follow the…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Xiang Wang , Zhifei Zhang , He Zhang , Zhe Lin , Yuqian Zhou , Qing Liu , Shiwei Zhang , Yijun Li , Shaoteng Liu , Haitian Zheng , Jason Kuen , Yuehuan Wang , Changxin Gao , Nong Sang

In medical imaging, 4D MRI enables dynamic 3D visualization, yet the trade-off between spatial and temporal resolution requires prolonged scan time that can compromise temporal fidelity--especially during rapid, large-amplitude motion.…

图像与视频处理 · 电气工程与系统科学 2025-06-10 Xuanru Zhou , Jiarun Liu , Shoujun Yu , Hao Yang , Cheng Li , Tao Tan , Shanshan Wang

Audio-visual deepfake detection typically employs a complementary multi-modal model to check the forgery traces in the video. These methods primarily extract forgery traces through audio-visual alignment, which results from the…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Fangda Wei , Miao Liu , Yingxue Wang , Jing Wang , Shenghui Zhao , Nan Li