English
Related papers

Related papers: Open-set Cross Modal Generalization via Multimodal…

200 papers

Learning joint representations across multiple modalities remains a central challenge in multimodal machine learning. Prevailing approaches predominantly operate in pairwise settings, aligning two modalities at a time. While some recent…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Stefanos Koutoupis , Michaela Areti Zervou , Konstantinos Kontras , Maarten De Vos , Panagiotis Tsakalides , Grigorios Tsagkatakis

Cross-modal contrastive distillation has recently been explored for learning effective 3D representations. However, existing methods focus primarily on modality-shared features, neglecting the modality-specific features during the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Yifan Zhang , Junhui Hou

Domain Generalization (DG) aims to enhance model robustness in unseen or distributionally shifted target domains through training exclusively on source domains. Although existing DG techniques, such as data manipulation, learning…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Hai Huang , Yan Xia , Sashuai Zhou , Hanting Wang , Shulei Wang , Zhou Zhao

Multi-modal Contrastive Representation learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across various modalities.…

Machine Learning · Computer Science 2023-10-20 Zehan Wang , Yang Zhao , Xize Cheng , Haifeng Huang , Jiageng Liu , Li Tang , Linjun Li , Yongqi Wang , Aoxiong Yin , Ziang Zhang , Zhou Zhao

The challenge of open-vocabulary recognition lies in the model has no clue of new categories it is applied to. Existing works have proposed different methods to embed category cues into the model, \eg, through few-shot fine-tuning,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Zehong Ma , Shiliang Zhang , Longhui Wei , Qi Tian

Multimodal representation learning is a challenging task in which previous work mostly focus on either uni-modality pre-training or cross-modality fusion. In fact, we regard modeling multimodal representation as building a skyscraper, where…

Computation and Language · Computer Science 2024-08-15 Ronghao Lin , Haifeng Hu

We propose to build omni-modal intelligence, which is capable of understanding any modality and learning universal representations. In specific, we propose a scalable pretraining paradigm, named Multimodal Context (MiCo), which can scale up…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Yiyuan Zhang , Handong Li , Jing Liu , Xiangyu Yue

Omni Large Language Models (Omni-LLMs) have demonstrated impressive capabilities in holistic multi-modal perception, yet they consistently falter in complex scenarios requiring synergistic omni-modal reasoning. Beyond understanding global…

Computation and Language · Computer Science 2026-04-08 Hongcheng Liu , Yuhao Wang , Zhe Chen , Pingjie Wang , Zhiyuan Zhu , Yixuan Hou , Yanfeng Wang , Yu Wang

Multimodal affective computing aims to predict humans' sentiment, emotion, intention, and opinion using language, acoustic, and visual modalities. However, current models often learn spurious correlations that harm generalization under…

Machine Learning · Computer Science 2026-04-21 Sijie Mai , Shiqin Han

Image clustering, which involves grouping images into different clusters without labels, is a key task in unsupervised learning. Although previous deep clustering methods have achieved remarkable results, they only explore the intrinsic…

Computer Vision and Pattern Recognition · Computer Science 2024-09-23 Haixin Zhang , Yongjun Li , Dong Huang

Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and textual modalities, allowing them to process image-text…

Computation and Language · Computer Science 2025-09-29 Xiaolong Wang , Zhaolu Kang , Wangyuxuan Zhai , Xinyue Lou , Yunghwei Lai , Ziyue Wang , Yawen Wang , Kaiyu Huang , Yile Wang , Peng Li , Yang Liu

Most multi-view clustering methods are limited by shallow models without sound nonlinear information perception capability, or fail to effectively exploit complementary information hidden in different views. To tackle these issues, we…

Machine Learning · Computer Science 2022-10-14 Fu Lele , Zhang Lei , Yang Jinghua , Chen Chuan , Zhang Chuanfu , Zheng Zibin

Multi-view subspace clustering aims to discover the hidden subspace structures from multiple views for robust clustering, and has been attracting considerable attention in recent years. Despite significant progress, most of the previous…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Xiaosha Cai , Dong Huang , Guang-Yu Zhang , Chang-Dong Wang

Technological advances facilitate the ability to acquire multimodal data, posing a challenge for recognition systems while also providing an opportunity to use the heterogeneous nature of the information to increase the generalization…

Machine Learning · Computer Science 2024-08-06 Paweł Zyblewski , Leandro L. Minku

Multimodal semantic segmentation enhances model robustness by exploiting cross-modal complementarities. However, existing methods often suffer from imbalanced modal dependencies, where overall performance degrades significantly once a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Jiaqi Tan , Xu Zheng , Fangyu Li , Yang Liu

Learning high-quality multi-modal entity representations is an important goal of multi-modal knowledge graph (MMKG) representation learning, which can enhance reasoning tasks within the MMKGs, such as MMKG completion (MMKGC). The main…

Artificial Intelligence · Computer Science 2025-04-08 Yichi Zhang , Zhuo Chen , Lingbing Guo , Yajing Xu , Binbin Hu , Ziqi Liu , Wen Zhang , Huajun Chen

To extend the reinforcement learning post-training paradigm to omni-modal models for concurrently bolstering video-audio understanding and collaborative reasoning, we propose OmniJigsaw, a generic self-supervised framework built upon a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Yiduo Jia , Muzhi Zhu , Hao Zhong , Mingyu Liu , Yuling Xi , Hao Chen , Bin Qin , Yongjie Yang , Zhenbo Luo , Chunhua Shen

Optical motion capture is a foundational technology driving advancements in cutting-edge fields such as virtual reality and film production. However, system performance suffers severely under large-scale marker occlusions common in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Chen Qian , Danyang Li , Xinran Yu , Zheng Yang , Qiang Ma

Commonsense reasoning in multimodal contexts remains a foundational challenge in artificial intelligence. We introduce Multimodal UNcommonsense(MUN), a benchmark designed to evaluate models' ability to handle scenarios that deviate from…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Yejin Son , Saejin Kim , Dongjun Min , Younjae Yu

In this thesis, we address the challenging problem of unpaired multi-view clustering (UMC), which aims to achieve effective joint clustering using unpaired samples observed across multiple views. Traditional incomplete multi-view clustering…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Like Xin , Wanqi Yang , Lei Wang , Ming Yang