English
Related papers

Related papers: Towards Flexible, Scalable, and Adaptive Multi-Mod…

200 papers

Audio-visual deepfake detection scrutinizes manipulations in public video using complementary multimodal cues. Current methods, which train on fused multimodal data for multimodal targets face challenges due to uncertainties and…

Multimedia · Computer Science 2024-01-12 Heqing Zou , Meng Shen , Yuchen Hu , Chen Chen , Eng Siong Chng , Deepu Rajan

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. This stems from insufficient cross-modal interaction and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Jiehui Huang , Yuechen Zhang , Xu He , Yuan Gao , Zhi Cen , Bin Xia , Yan Zhou , Xin Tao , Pengfei Wan , Jiaya Jia

Highly accurate datasets from numerical or physical experiments are often expensive and time-consuming to acquire, posing a significant challenge for applications that require precise evaluations, potentially across multiple scenarios and…

Machine Learning · Computer Science 2026-02-06 Paolo Conti , Mengwu Guo , Attilio Frangi , Andrea Manzoni

Learning to reliably perceive and understand the scene is an integral enabler for robots to operate in the real-world. This problem is inherently challenging due to the multitude of object types as well as appearance changes caused by…

Computer Vision and Pattern Recognition · Computer Science 2021-11-05 Abhinav Valada , Rohit Mohan , Wolfram Burgard

Most methods for conditional video synthesis use a single modality as the condition. This comes with major limitations. For example, it is problematic for a model conditioned on an image to generate a specific motion trajectory desired by…

Computer Vision and Pattern Recognition · Computer Science 2022-03-08 Ligong Han , Jian Ren , Hsin-Ying Lee , Francesco Barbieri , Kyle Olszewski , Shervin Minaee , Dimitris Metaxas , Sergey Tulyakov

With the increasing availability of diverse data types, particularly images and time series data from medical experiments, there is a growing demand for techniques designed to combine various modalities of data effectively. Our motivation…

Image and Video Processing · Electrical Eng. & Systems 2024-05-27 Ali Rasekh , Reza Heidari , Amir Hosein Haji Mohammad Rezaie , Parsa Sharifi Sedeh , Zahra Ahmadi , Prasenjit Mitra , Wolfgang Nejdl

Synthesizing MR imaging sequences is highly relevant in clinical practice, as single sequences are often missing or are of poor quality (e.g. due to motion). Naturally, the idea arises that a target modality would benefit from multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Hongwei Li , Johannes C. Paetzold , Anjany Sekuboyina , Florian Kofler , Jianguo Zhang , Jan S. Kirschke , Benedikt Wiestler , Bjoern Menze

Learning multi-modal representations is an essential step towards real-world robotic applications, and various multi-modal fusion models have been developed for this purpose. However, we observe that existing models, whose objectives are…

Machine Learning · Computer Science 2021-06-22 Chenzhuang Du , Tingle Li , Yichen Liu , Zixin Wen , Tianyu Hua , Yue Wang , Hang Zhao

Self-supervised learning is an efficient pre-training method for medical image analysis. However, current research is mostly confined to specific-modality data pre-training, consuming considerable time and resources without achieving…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Yiwen Ye , Yutong Xie , Jianpeng Zhang , Ziyang Chen , Qi Wu , Yong Xia

Diffusion models arise as a powerful generative tool recently. Despite the great progress, existing diffusion models mainly focus on uni-modal control, i.e., the diffusion process is driven by only one modality of condition. To further…

Computer Vision and Pattern Recognition · Computer Science 2023-04-21 Ziqi Huang , Kelvin C. K. Chan , Yuming Jiang , Ziwei Liu

Recent multimodal face generation models address the spatial control limitations of text-to-image diffusion models by augmenting text-based conditioning with spatial priors such as segmentation masks, sketches, or edge maps. This multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Bharath Krishnamurthy , Ajita Rattani

In this work, we introduce a new approach for face stylization. Despite existing methods achieving impressive results in this task, there is still room for improvement in generating high-quality artistic faces with diverse styles and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Mengtian Li , Yi Dong , Minxuan Lin , Haibin Huang , Pengfei Wan , Chongyang Ma

Decoder-only discrete-token language models have recently achieved significant success in automatic speech recognition. However, systematic analyses of how different modalities impact performance in specific scenarios remain limited. In…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Yiwen Guan , Viet Anh Trinh , Vivek Voleti , Jacob Whitehill

Regional facial image synthesis conditioned on semantic mask has achieved great success using generative adversarial networks. However, the appearance of different regions may be inconsistent with each other when conducting regional image…

Multimedia · Computer Science 2021-04-30 Cong Wang , Fan Tang , Yong Zhang , Weiming Dong , Tieru Wu

With the growing success of multi-modal learning, research on the robustness of multi-modal models, especially when facing situations with missing modalities, is receiving increased attention. Nevertheless, previous studies in this domain…

Artificial Intelligence · Computer Science 2023-10-11 Siting Li , Chenzhuang Du , Yue Zhao , Yu Huang , Hang Zhao

Multimodal MR image synthesis aims to generate missing modality images by effectively fusing and mapping from a subset of available MRI modalities. Most existing methods adopt an image-to-image translation paradigm, treating multiple…

Image and Video Processing · Electrical Eng. & Systems 2025-04-29 Tao Song , Yicheng Wu , Minhao Hu , Xiangde Luo , Linda Wei , Guotai Wang , Yi Guo , Feng Xu , Shaoting Zhang

Multi-modality imaging improves disease diagnosis and reveals distinct deviations in tissues with anatomical properties. The existence of completely aligned and paired multi-modality neuroimaging data has proved its effectiveness in brain…

Image and Video Processing · Electrical Eng. & Systems 2023-09-26 Guoyang Xie , Yawen Huang , Jinbao Wang , Jiayi Lyu , Feng Zheng , Yefeng Zheng , Yaochu Jin

Fusing and balancing multi-modal inputs from novel sensors for dense prediction tasks, particularly semantic segmentation, is critically important yet remains a significant challenge. One major limitation is the tendency of multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Xu Zheng , Yuanhuiyi Lyu , Lutao Jiang , Danda Pani Paudel , Luc Van Gool , Xuming Hu

We present UniRef-Image-Edit, a high-performance multi-modal generation system that unifies single-image editing and multi-image composition within a single framework. Existing diffusion-based editing methods often struggle to maintain…

Human communication is multi-modal; e.g., face-to-face interaction involves auditory signals (speech) and visual signals (face movements and hand gestures). Hence, it is essential to exploit multiple modalities when designing machine…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Marah Halawa , Florian Blume , Pia Bideau , Martin Maier , Rasha Abdel Rahman , Olaf Hellwich