English
Related papers

Related papers: Towards Flexible, Scalable, and Adaptive Multi-Mod…

200 papers

Recent advances in semantic segmentation of multi-modal remote sensing images have significantly improved the accuracy of tree cover mapping, supporting applications in urban planning, forest monitoring, and ecological assessment.…

Image and Video Processing · Electrical Eng. & Systems 2025-12-16 Yuanyuan Gui , Wei Li , Yinjian Wang , Xiang-Gen Xia , Mauro Marty , Christian Ginzler , Zuyuan Wang

Recent studies have focused on utilizing multi-modal data to develop robust models for facial Action Unit (AU) detection. However, the heterogeneity of multi-modal data poses challenges in learning effective representations. One such…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Xiang Zhang , Huiyuan Yang , Taoyue Wang , Xiaotian Li , Lijun Yin

Electroencephalography (EEG)-based multimodal learning integrates brain signals with complementary modalities to improve mental state assessment, providing great clinical potential. The effectiveness of such paradigms largely depends on the…

Machine Learning · Computer Science 2026-05-12 Runhe Zhou , Shanglin Li , Guanxiang Huang , Xinliang Zhou , Qibin Zhao , Motoaki Kawanabe , Yi Ding , Cuntai Guan

As synthetic media, including video, audio, and text, become increasingly indistinguishable from real content, the risks of misinformation, identity fraud, and social manipulation escalate. This survey traces the evolution of deepfake…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Ping Liu , Qiqi Tao , Joey Tianyi Zhou

Longitudinal face recognition in children remains challenging due to rapid and nonlinear facial growth, which causes template drift and increasing verification errors over time. This work investigates whether synthetic face data can act as…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Afzal Hossain , Stephanie Schuckers

Cross-modality synthesis (CMS), super-resolution (SR), and their combination (CMSR) have been extensively studied for magnetic resonance imaging (MRI). Their primary goals are to enhance the imaging quality by synthesizing the desired…

Image and Video Processing · Electrical Eng. & Systems 2023-11-15 Zhiyun Song , Zengxin Qi , Xin Wang , Xiangyu Zhao , Zhenrong Shen , Sheng Wang , Manman Fei , Zhe Wang , Di Zang , Dongdong Chen , Linlin Yao , Qian Wang , Xuehai Wu , Lichi Zhang

Multi-modal entity alignment aims to identify equivalent entities between two multi-modal Knowledge graphs by integrating multi-modal data, such as images and text, to enrich the semantic representations of entities. However, existing…

Artificial Intelligence · Computer Science 2026-01-21 Zhifei Li , Ziyue Qin , Xiangyu Luo , Xiaoju Hou , Yue Zhao , Miao Zhang , Zhifang Huang , Kui Xiao , Bing Yang

To detect bias in face recognition networks, it can be useful to probe a network under test using samples in which only specific attributes vary in some controlled way. However, capturing a sufficiently large dataset with specific control…

Computer Vision and Pattern Recognition · Computer Science 2020-12-11 Nataniel Ruiz , Barry-John Theobald , Anurag Ranjan , Ahmed Hussein Abdelaziz , Nicholas Apostoloff

As information exists in various modalities in real world, effective interaction and fusion among multimodal information plays a key role for the creation and perception of multimodal data in computer vision and deep learning research. With…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Fangneng Zhan , Yingchen Yu , Rongliang Wu , Jiahui Zhang , Shijian Lu , Lingjie Liu , Adam Kortylewski , Christian Theobalt , Eric Xing

With rapid advancements in generative modeling, deepfake techniques are increasingly narrowing the gap between real and synthetic videos, raising serious privacy and security concerns. Beyond traditional face swapping and reenactment, an…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Tharun Anand , Siva Sankar Sajeev , Pravin Nair

Cross-modal image synthesis is a topical problem in medical image computing. Existing methods for image synthesis are either tailored to a specific application, require large scale training sets, or are based on partitioning images into…

Computer Vision and Pattern Recognition · Computer Science 2017-06-16 Yawen Huang , Ling Shao , Alejandro F. Frangi

Latent space-based facial attribute editing methods have gained popularity in applications such as digital entertainment, virtual avatar creation, and human-computer interaction systems due to their potential for efficient and flexible…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Bo Liu , Xuan Cui , Run Zeng , Wei Duan , Chongwen Liu , Jinrui Qian , Lianggui Tang , Hongping Gan

Multimodal learning, which integrates data from diverse sensory modes, plays a pivotal role in artificial intelligence. However, existing multimodal learning methods often struggle with challenges where some modalities appear more dominant…

Machine Learning · Computer Science 2024-04-02 Xiaohui Zhang , Jaehong Yoon , Mohit Bansal , Huaxiu Yao

Previous deepfake detection methods mostly depend on low-level textural features vulnerable to perturbations and fall short of detecting unseen forgery methods. In contrast, high-level semantic features are less susceptible to perturbations…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Ziyuan Fang , Hanqing Zhao , Tianyi Wei , Wenbo Zhou , Ming Wan , Zhanyi Wang , Weiming Zhang , Nenghai Yu

Purpose: Different Magnetic resonance imaging (MRI) modalities of the same anatomical structure are required to present different pathological information from the physical level for diagnostic needs. However, it is often difficult to…

Image and Video Processing · Electrical Eng. & Systems 2021-09-15 Yuchen Fei , Bo Zhan , Mei Hong , Xi Wu , Jiliu Zhou , Yan Wang

Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based multimodal fusion methods ignore the inter-modality…

Computer Vision and Pattern Recognition · Computer Science 2023-05-17 Heqing Zou , Meng Shen , Chen Chen , Yuchen Hu , Deepu Rajan , Eng Siong Chng

Although humans engaged in face-to-face conversation simultaneously communicate both verbally and non-verbally, methods for joint and unified synthesis of speech audio and co-speech 3D gesture motion from text are a new and emerging field.…

Human-Computer Interaction · Computer Science 2024-05-01 Shivam Mehta , Anna Deichler , Jim O'Regan , Birger Moëll , Jonas Beskow , Gustav Eje Henter , Simon Alexanderson

While the accuracy of face recognition systems has improved significantly in recent years, the datasets used to train these models are often collected through web crawling without the explicit consent of users, raising ethical and privacy…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Anjith George , Sebastien Marcel

As medical diagnoses increasingly leverage multimodal data, machine learning models are expected to effectively fuse heterogeneous information while remaining robust to missing modalities. In this work, we propose a novel multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Yi Gu , Kuniaki Saito , Jiaxin Ma

Human face synthesis and manipulation are increasingly important in entertainment and AI, with a growing demand for highly realistic, identity-preserving images even when only unpaired, unaligned datasets are available. We study unpaired…

Machine Learning · Computer Science 2026-01-05 Collin Guo , Yi Qian