中文
相关论文

相关论文: Leveraging Generic Foundation Models for Multimoda…

200 篇论文

In this work, we introduce the Multiple Embedding Model for EHR (MEME), an approach that serializes multimodal EHR tabular data into text using pseudo-notes, mimicking clinical text generation. This conversion not only preserves better…

计算与语言 · 计算机科学 2025-07-08 Simon A. Lee , Sujay Jain , Alex Chen , Kyoka Ono , Jennifer Fang , Akos Rudas , Jeffrey N. Chiang

Cloud segmentation is a critical challenge in remote sensing image interpretation, as its accuracy directly impacts the effectiveness of subsequent data processing and analysis. Recently, vision foundation models (VFM) have demonstrated…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Xuechao Zou , Shun Zhang , Kai Li , Shiying Wang , Junliang Xing , Lei Jin , Congyan Lang , Pin Tao

The rapid advancement of autonomous systems, including self-driving vehicles and drones, has intensified the need to forge true Spatial Intelligence from multi-modal onboard sensor data. While foundation models excel in single-modal…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Song Wang , Lingdong Kong , Xiaolu Liu , Hao Shi , Wentong Li , Jianke Zhu , Steven C. H. Hoi

Multi-modal foundation models are typically trained on millions of pairs of natural images and text captions, frequently obtained through web-crawling approaches. Although such models depict excellent generative capabilities, they do not…

计算机视觉与模式识别 · 计算机科学 2023-01-03 Pierre Chambon , Christian Bluethgen , Curtis P. Langlotz , Akshay Chaudhari

Domain adaptation is crucial for transferring the knowledge from the source labeled CT dataset to the target unlabeled MR dataset in abdominal multi-organ segmentation. Meanwhile, it is highly desirable to avoid the high annotation cost…

计算机视觉与模式识别 · 计算机科学 2022-05-31 Jin Hong , Yu-Dong Zhang , Weitian Chen

Multimodal embedding models, built upon causal Vision Language Models (VLMs), have shown promise in various tasks. However, current approaches face three key limitations: the use of causal attention in VLM backbones is suboptimal for…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Haonan Chen , Hong Liu , Yuping Luo , Liang Wang , Nan Yang , Furu Wei , Zhicheng Dou

We introduce Neural Organ Transplantation (NOT), a modular adaptation framework that enables trained transformer layers to function as reusable transferable checkpoints for domain adaptation. Unlike conventional fine-tuning approaches that…

机器学习 · 计算机科学 2026-01-21 Ahmad Al-Zuraiqi

Visual model-based reinforcement learning (MBRL) agents can perform well on the training distribution, but often break down once the test environment shifts. In visual MBRL, recognizing that a shift has occurred is often the easier part;…

机器学习 · 计算机科学 2026-05-01 Haiyang Zhao

Multimodal Emotion Recognition in Conversations remains a challenging task due to the complex interplay of textual, acoustic and visual signals. While recent models have improved performance via advanced fusion strategies, they often lack…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Guanyu Hu , Dimitrios Kollias , Xinyu Yang

We survey applications of pretrained foundation models in robotics. Traditional deep learning models in robotics are trained on small datasets tailored for specific tasks, which limits their adaptability across diverse applications. In…

State-of-the-art large language and vision models are trained over trillions of tokens that are aggregated from a large variety of sources. As training data collections grow, manually managing the samples becomes time-consuming, tedious,…

机器学习 · 计算机科学 2026-02-03 Maximilian Böther , Xiaozhe Yao , Tolga Kerimoglu , Dan Graur , Viktor Gsteiger , Ana Klimovic

Foundation models pre-trained on web-scale data are shown to encapsulate extensive world knowledge beneficial for robotic manipulation in the form of task planning. However, the actual physical implementation of these plans often relies on…

机器人学 · 计算机科学 2024-03-14 Haoxu Huang , Fanqi Lin , Yingdong Hu , Shengjie Wang , Yang Gao

Foundation models have demonstrated remarkable generalization, data efficiency, and robustness properties across various domains. In this paper, we explore the feasibility of foundation models for applications in the control domain. The…

机器学习 · 计算机科学 2024-12-18 Martin Ziegler , Andres Felipe Posada-Moreno , Friedrich Solowjow , Sebastian Trimpe

Developing robust and versatile deep-learning models is essential for enhancing diagnostic accuracy and guiding clinical interventions in medical imaging, but it requires a large amount of annotated data. The advancement of deep learning…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Nahid Ul Islam , DongAo Ma , Jiaxuan Pang , Shivasakthi Senthil Velan , Michael Gotway , Jianming Liang

Evaluating foundation models under appropriate adaptation settings is essential for understanding the quality and transferability of the learned representations. Recent EEG foundation models have demonstrated promising transfer capabilities…

机器学习 · 计算机科学 2026-05-28 Aditya Kommineni , Emily Zhou , Kleanthis Avramidis , Tiantian Feng , Shrikanth Narayanan

Graph Neural Networks (GNNs) have shown promise in learning dynamic functional connectivity for distinguishing phenotypes from human brain networks. However, obtaining extensive labeled clinical data for training is often…

机器学习 · 计算机科学 2025-05-06 Jungwon Choi , Hyungi Lee , Byung-Hoon Kim , Juho Lee

Earth observation (EO) in open-world settings presents a unique challenge: different applications rely on diverse sensor modalities, each with varying ground sampling distances, spectral ranges, and numbers of spectral bands. However,…

The present study explores the interpretability of latent spaces produced by time series foundation models, focusing on their potential for visual analysis tasks. Specifically, we evaluate the MOMENT family of models, a set of…

Although large-scale visual foundation models (VFMs) achieve remarkable performance in semantic understanding, they still underperform in instance-aware dense prediction tasks. They exhibit different biases in representation: for instance,…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Yachan Guo , JoseLuis Gomez Zurita , Danna Xue , Yi Xiao , AntonioManuel Lopez Pena

This paper investigates an under-explored but important problem: given a collection of pre-trained neural networks, predicting their performance on each multi-modal task without fine-tuning them, such as image recognition, referring,…

机器学习 · 计算机科学 2023-08-14 Fanqing Meng , Wenqi Shao , Zhanglin Peng , Chonghe Jiang , Kaipeng Zhang , Yu Qiao , Ping Luo