中文
相关论文

相关论文: Leveraging Generic Foundation Models for Multimoda…

200 篇论文

With the rapid advancement of large language models (LLMs), foundational models (FMs) have seen significant advancements. Healthcare is one of the most crucial application areas for these FMs, given the significant time and effort required…

计算机视觉与模式识别 · 计算机科学 2024-11-18 Kaito Baba , Ryota Yagi , Junichiro Takahashi , Risa Kishikawa , Satoshi Kodera

Leveraging pre-trained visual language models has become a widely adopted approach for improving performance in downstream visual question answering (VQA) applications. However, in the specialized field of medical VQA, the scarcity of…

计算机视觉与模式识别 · 计算机科学 2024-03-06 Gang Liu , Hongyang Li , Zerui He , Shenjun Zhong

Foundational models are trained on extensive datasets to capture the general trends of a domain. However, in medical imaging, the scarcity of data makes pre-training for every domain, modality, or task challenging. Continual learning offers…

图像与视频处理 · 电气工程与系统科学 2025-08-20 Mohammad Areeb Qazi , Munachiso S Nwadike , Ibrahim Almakky , Mohammad Yaqub , Numan Saeed

We present Video Pre-trained Transformer. VPT uses four SOTA encoder models from prior work to convert a video into a sequence of compact embeddings. Our backbone, based on a reference Flan-T5-11B architecture, learns a universal…

计算机视觉与模式识别 · 计算机科学 2023-04-21 Kastan Day , Daniel Christl , Rohan Salvi , Pranav Sriram

The recent boom of large pre-trained models witnesses remarkable success in developing foundation models (FMs) for time series forecasting. Despite impressive performance across diverse downstream forecasting tasks, existing time series FMs…

机器学习 · 计算机科学 2025-10-23 Hui He , Kun Yi , Yuanchi Ma , Qi Zhang , Zhendong Niu , Guansong Pang

The demand for high-resolution subsurface imaging and continuous Earth monitoring has driven rapid growth in active and passive seismic data from dense geophone deployments, distributed acoustic sensing (DAS) arrays, and large-scale 2D and…

地球物理 · 物理学 2026-05-13 Jiahua Zhao , Umair bin Waheed , Jing Sun , Yang Cui , Nikos Savva , Eric Verschuur

Video Joint Embedding Predictive Architectures (V-JEPA) learn generalizable off-the-shelf video representation by predicting masked regions in latent space with an exponential moving average (EMA)-updated teacher. While EMA prevents…

机器学习 · 计算机科学 2025-09-30 Xianhang Li , Chen Huang , Chun-Liang Li , Eran Malach , Josh Susskind , Vimal Thilak , Etai Littwin

Purpose: Applying pre-trained medical deep learning segmentation models on out-of-domain images often yields predictions of insufficient quality. In this study, we propose to use a powerful generalizing descriptor along with augmentation to…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Christian Weihsbach , Christian N. Kruse , Alexander Bigalke , Mattias P. Heinrich

Foundation models trained on electronic health records show strong performance on many clinical prediction tasks but are limited by sparse and irregular documentation. Wearable devices provide dense continuous physiological signals but lack…

机器学习 · 计算机科学 2026-01-21 Yuanyun Zhang , Han Zhou , Li Feng , Yilin Hong , Shi Li

Foundation models (FMs) have emerged as a transformative paradigm in medical image analysis, offering the potential to provide generalizable, task-agnostic solutions across a wide range of clinical tasks and imaging modalities. Their…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Karma Phuntsho , Abdullah , Kyungmi Lee , Ickjai Lee , Euijoon Ahn

State-of-the-art vessel segmentation methods typically require large-scale annotated datasets and suffer from severe performance degradation under domain shifts. In clinical practice, however, acquiring extensive annotations for every new…

图像与视频处理 · 电气工程与系统科学 2026-03-02 Kirato Yoshihara , Yohei Sugawara , Yuta Tokuoka , Lihang Hong

Medical vision foundation models remain limited in downstream tasks, particularly volumetric medical image segmentation. While fine-tuning on labeled target-domain data improves performance, existing approaches typically rely on randomly…

图像与视频处理 · 电气工程与系统科学 2026-05-07 Jin Yang , Daniel S. Marcus , Aristeidis Sotiras

Medical image grounding aims to align natural language phrases with specific regions in medical images, serving as a foundational task for intelligent diagnosis, visual question answering (VQA), and automated report generation (MRG).…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Ziye Deng , Ruihan He , Jiaxiang Liu , Yuan Wang , Zijie Meng , Songtao Jiang , Yong Xie , Zuozhu Liu

Multivariate time series underpin modern critical infrastructure, making the prediction of anomalies a vital necessity for proactive risk mitigation. While Joint-Embedding Predictive Architectures (JEPA) offer a promising framework for…

机器学习 · 计算机科学 2026-02-05 Yanan He , Yunshi Wen , Xin Wang , Tengfei Ma

Prediction tasks in digital pathology are challenging due to the massive size of whole-slide images (WSIs) and the weak nature of training signals. Advances in computing, data availability, and self-supervised learning (SSL) have paved the…

图像与视频处理 · 电气工程与系统科学 2026-02-02 Vishwesh Ramanathan , Tony Xu , Pushpak Pati , Faruk Ahmed , Maged Goubran , Anne L. Martel

This article discusses the opportunities, applications and future directions of large-scale pre-trained models, i.e., foundation models, for analyzing medical images. Medical foundation models have immense potential in solving a wide range…

图像与视频处理 · 电气工程与系统科学 2023-11-23 Shaoting Zhang , Dimitris Metaxas

Federated Learning with LoRA fine-tuning offers an efficient and privacy-aware solution for institutions to collaboratively leverage their large datasets to train VLLMs. However, participating institutions often possess heterogeneous…

机器学习 · 计算机科学 2026-05-19 Lishan Yang , Wei Emma Zhang , Nam Kha Nguygen , Po Hu , Yanjun Shu , Weitong Chen , Mong Yuan Sim

Vision-Language Foundation Models (VLFMs) exhibit remarkable generalization, yet their direct application to medical ultrasound is severely hindered by a profound modality gap. The unique acoustic physics of ultrasound, characterized by…

Improving the generalization capabilities of general-purpose robotic manipulation agents in the real world has long been a significant challenge. Existing approaches often rely on collecting large-scale robotic data which is costly and…

机器人学 · 计算机科学 2025-02-10 Jiange Yang , Wenhui Tan , Chuhao Jin , Keling Yao , Bei Liu , Jianlong Fu , Ruihua Song , Gangshan Wu , Limin Wang

Automatic surgical activity recognition enables more intelligent surgical devices and a more efficient workflow. Integration of such technology in new operating rooms has the potential to improve care delivery to patients and decrease…

计算机视觉与模式识别 · 计算机科学 2022-07-08 Ali Mottaghi , Aidean Sharghi , Serena Yeung , Omid Mohareri