English
Related papers

Related papers: LANISTR: Multimodal Learning from Structured and U…

200 papers

Data is one of the essential ingredients to power deep learning research. Small datasets, especially specific to medical institutes, bring challenges to deep learning training stage. This work aims to develop a practical deep multimodal…

Machine Learning · Computer Science 2019-02-26 Faik Aydin , Maggie Zhang , Michelle Ananda-Rajah , Gholamreza Haffari

Building generalizable medical AI systems requires pretraining strategies that are data-efficient and domain-aware. Unlike internet-scale corpora, clinical datasets such as MIMIC-CXR offer limited image counts and scarce annotations, but…

State-of-the-art computer vision models are mostly trained with supervised learning using human-labeled images, which limits their scalability due to the expensive annotation cost. While self-supervised representation learning has achieved…

Computer Vision and Pattern Recognition · Computer Science 2023-03-13 Junnan Li , Silvio Savarese , Steven C. H. Hoi

Different modalities hold considerable gaps in optimization trajectories, including speeds and paths, which lead to modality laziness and modality clash when jointly training multimodal models, resulting in insufficient and imbalanced…

Machine Learning · Computer Science 2025-06-17 Xiaoyu Ma , Hao Chen , Yongjian Deng

Machine learning holds promise for advancing clinical decision support, yet it remains unclear when multimodal learning truly helps in practice, particularly under modality missingness and fairness constraints. In this work, we conduct a…

Machine Learning · Computer Science 2026-03-02 Kejing Yin , Haizhou Xu , Wenfang Yao , Chen Liu , Zijie Chen , Yui Haang Cheung , William K. Cheung , Jing Qin

Multi-modal contrastive learning with language supervision has presented a paradigm shift in modern machine learning. By pre-training on a web-scale dataset, multi-modal contrastive learning can learn high-quality representations that…

Machine Learning · Computer Science 2024-11-06 Wei Huang , Andi Han , Yongqiang Chen , Yuan Cao , Zhiqiang Xu , Taiji Suzuki

Data-efficient learning aims to eliminate redundancy in large training datasets by training models on smaller subsets of the most informative examples. While data selection has been extensively explored for vision models and large language…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Nilay Naharas , Dang Nguyen , Nesihan Bulut , Mohammadhossein Bateni , Vahab Mirrokni , Baharan Mirzasoleiman

Multimodal AI has demonstrated superior performance over unimodal approaches by leveraging diverse data sources for more comprehensive analysis. However, applying this effectiveness in healthcare is challenging due to the limited…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Pranav Poudel , Prashant Shrestha , Sanskar Amgain , Yash Raj Shrestha , Prashnna Gyawali , Binod Bhattarai

Vision-Language Models (VLMs) have achieved remarkable progress in multimodal understanding, yet their positional encoding mechanisms remain suboptimal. Existing approaches uniformly assign positional indices to all tokens, overlooking…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Ruoxiang Huang , Zhen Yuan

Fusing multi-modal data can improve the performance of deep learning models. However, missing modalities are common for medical data due to patients' specificity, which is detrimental to the performance of multi-modal models in…

Image and Video Processing · Electrical Eng. & Systems 2023-09-28 Muyu Wang , Shiyu Fan , Yichen Li , Hui Chen

Vision-language models (VLMs) have made significant strides in cross-modal understanding through large-scale paired datasets. However, in fashion domain, datasets often exhibit a disparity between the information conveyed in image and text.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Chull Hwan Song , Taebaek Hwang , Jooyoung Yoon , Shunghyun Choi , Yeong Hyeon Gu

Due to the ever-growing diversity of the data source, multi-modality feature learning has attracted more and more attention. However, most of these methods are designed by jointly learning feature representation from multi-modalities that…

Computer Vision and Pattern Recognition · Computer Science 2020-06-09 Danfeng Hong , Jocelyn Chanussot , Naoto Yokoya , Jian Kang , Xiao Xiang Zhu

This paper presents MAST, a new model for Multimodal Abstractive Text Summarization that utilizes information from all three modalities -- text, audio and video -- in a multimodal video. Prior work on multimodal abstractive text…

Computation and Language · Computer Science 2020-10-19 Aman Khullar , Udit Arora

In this paper, we study a novel problem in egocentric action recognition, which we term as "Multimodal Generalization" (MMG). MMG aims to study how systems can generalize when data from certain modalities is limited or even completely…

Computer Vision and Pattern Recognition · Computer Science 2023-05-15 Xinyu Gong , Sreyas Mohan , Naina Dhingra , Jean-Charles Bazin , Yilei Li , Zhangyang Wang , Rakesh Ranjan

Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require pre-training on a large set of parallel image-text data,…

Computer Vision and Pattern Recognition · Computer Science 2022-03-02 Mingyang Zhou , Licheng Yu , Amanpreet Singh , Mengjiao Wang , Zhou Yu , Ning Zhang

Unsupervised methods have proven effective for discriminative tasks in a single-modality scenario. In this paper, we present a multimodal framework for learning sparse representations that can capture semantic correlation between…

Machine Learning · Computer Science 2016-03-03 Miriam Cha , Youngjune Gwon , H. T. Kung

Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical problem for educational support and skill transfer. Yet such face-to-face instructional scenes have not…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Yuki Sakai , Ryosuke Furuta , Juichun Yen , Yoichi Sato

The increase in parameter size of multimodal large language models (MLLMs) introduces significant capabilities, particularly in-context learning, where MLLMs enhance task performance without updating pre-trained parameters. This…

Computation and Language · Computer Science 2024-11-13 Yang Luo , Zangwei Zheng , Zirui Zhu , Yang You

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Jingzhi Li , Changjiang Luo , Ruoyu Chen , Hua Zhang , Wenqi Ren , Jianhou Gan , Xiaochun Cao

As two important textual modalities in electronic health records (EHR), both structured data (clinical codes) and unstructured data (clinical narratives) have recently been increasingly applied to the healthcare domain. Most existing…

Computation and Language · Computer Science 2022-11-01 Sicen Liu , Xiaolong Wang , Yongshuai Hou , Ge Li , Hui Wang , Hui Xu , Yang Xiang , Buzhou Tang