English
Related papers

Related papers: M&M: Multimodal-Multitask Model Integrating Audiov…

200 papers

The new educational models such as smart learning environments use of digital and context-aware devices to facilitate the learning process. In this new educational scenario, a huge quantity of multimodal students' data from a variety of…

Computers and Society · Computer Science 2025-11-27 Wilson Chango , Juan A. Lara , Rebeca Cerezo , Cristóbal Romero

In this paper, we propose a novel architecture for multi-modal speech and text input. We combine pretrained speech and text encoders using multi-headed cross-modal attention and jointly fine-tune on the target problem. The resultant…

Computation and Language · Computer Science 2022-04-21 Karan Singla , Daniel Pressel , Ryan Price , Bhargav Srinivas Chinnari , Yeon-Jun Kim , Srinivas Bangalore

In driving scenarios, automobile active safety systems are increasingly incorporating deep learning technology. These systems typically need to handle multiple tasks simultaneously, such as detecting fatigue driving and recognizing the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Shulei Qu , Zhenguo Gao , Xiaowei Chen , Na Li , Yakai Wang , Xiaoxiao Wu

In this work, we introduce Context-Aware MultiModal Learner (CaMML), for tuning large multimodal models (LMMs). CaMML, a lightweight module, is crafted to seamlessly integrate multimodal contextual samples into large models, thereby…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Yixin Chen , Shuai Zhang , Boran Han , Tong He , Bo Li

Recent advances in multimodal question answering have primarily focused on combining heterogeneous modalities or fine-tuning multimodal large language models. While these approaches have shown strong performance, they often rely on a…

Computation and Language · Computer Science 2026-04-22 Krishna Singh Rajput , Tejas Anvekar , Chitta Baral , Vivek Gupta

Despite excelling on multimodal benchmarks, vision-language models (VLMs) largely remain a black box. In this paper, we propose a novel interpretability framework to systematically analyze the internal mechanisms of VLMs, focusing on the…

Artificial Intelligence · Computer Science 2025-12-12 Yanbei Jiang , Xueqi Ma , Shu Liu , Sarah Monazam Erfani , Tongliang Liu , James Bailey , Jey Han Lau , Krista A. Ehinger

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in interpreting visual layouts and text. However, a significant challenge remains in their ability to interpret robustly and reason over multi-tabular data presented as…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Anshul Singh , Chris Biemann , Jan Strich

Multi-view action recognition (MVAR) leverages complementary temporal information from different views to improve the learning performance. Obtaining informative view-specific representation plays an essential role in MVAR. Attention has…

Computer Vision and Pattern Recognition · Computer Science 2020-11-30 Yue Bai , Zhiqiang Tao , Lichen Wang , Sheng Li , Yu Yin , Yun Fu

Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In light of the recent advances in multimodal large language models (MLLMs), such as GPT-4v and LLaVA,…

Artificial Intelligence · Computer Science 2024-12-03 Jianyi Zhang , Hao Frank Yang , Ang Li , Xin Guo , Pu Wang , Haiming Wang , Yiran Chen , Hai Li

The significance of mental health classification is paramount in contemporary society, where digital platforms serve as crucial sources for monitoring individuals' well-being. However, existing social media mental health datasets primarily…

Computation and Language · Computer Science 2024-11-08 Rina Carines Cabral , Siwen Luo , Josiah Poon , Soyeon Caren Han

Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial factors: the model architecture to facilitate the feature…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Tianyu Yu , Jinyi Hu , Yuan Yao , Haoye Zhang , Yue Zhao , Chongyi Wang , Shan Wang , Yinxv Pan , Jiao Xue , Dahai Li , Zhiyuan Liu , Hai-Tao Zheng , Maosong Sun

In recent years, deep learning has shown near-expert performance in segmenting complex medical tissues and tumors. However, existing models are often task-specific, with performance varying across modalities and anatomical regions.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 T-Mai Bui , Fares Bougourzi , Fadi Dornaika , Vinh Truong Hoang

Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for multi-modal pretraining, which can benefit from modality…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Haiyang Xu , Qinghao Ye , Ming Yan , Yaya Shi , Jiabo Ye , Yuanhong Xu , Chenliang Li , Bin Bi , Qi Qian , Wei Wang , Guohai Xu , Ji Zhang , Songfang Huang , Fei Huang , Jingren Zhou

Medical visual question answering (Med-VQA) is a crucial multimodal task in clinical decision support and telemedicine. Recent self-attention based methods struggle to effectively handle cross-modal semantic alignments between vision and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Qiangguo Jin , Xianyao Zheng , Hui Cui , Changming Sun , Yuqi Fang , Cong Cong , Ran Su , Leyi Wei , Ping Xuan , Junbo Wang

Multimodal Large Language Models (MLLMs) have demonstrated impressive abilities across various tasks, including visual question answering and chart comprehension, yet existing benchmarks for chart-related tasks fall short in capturing the…

Computation and Language · Computer Science 2025-02-11 Zifeng Zhu , Mengzhao Jia , Zhihan Zhang , Lang Li , Meng Jiang

The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Pengcheng Zhao , Jinxing Zhou , Yang Zhao , Dan Guo , Yanxiang Chen

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can't large language models replicate this holistic understanding? Through a systematic analysis of existing training paradigms in…

Multi-modal learning has shown exceptional performance in various tasks, especially in medical applications, where it integrates diverse medical information for comprehensive diagnostic evidence. However, there still are several challenges…

Machine Learning · Computer Science 2024-11-19 Lin Fan , Yafei Ou , Cenyang Zheng , Pengyu Dai , Tamotsu Kamishima , Masayuki Ikebe , Kenji Suzuki , Xun Gong

Recent Multimodal Large Language Models (MLLMs) have demonstrated significant progress in perceiving and reasoning over multimodal inquiries, ushering in a new research era for foundation models. However, vision-language misalignment in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Wei-Yao Wang , Zhao Wang , Helen Suzuki , Yoshiyuki Kobayashi

Recent advances in Multimodal Large Language Models (MLLMs) have shown impressive reasoning capabilities across vision-language tasks, yet still face the challenge of compute-difficulty mismatch. Through empirical analyses, we identify that…

Machine Learning · Computer Science 2026-03-17 Huijie Guo , Jingyao Wang , Lingyu Si , Jiahuan Zhou , Changwen Zheng , Wenwen Qiang