English
Related papers

Related papers: Multimodal Foundation Models For Echocardiogram In…

200 papers

Multi-modal models require aligned, shared embedding spaces. However, common CLIP-based approaches need large amounts of samples and do not natively support 3D or tabular data, both of which are crucial in the medical domain. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Jakob Krogh Petersen , Valdemar Licht , Mads Nielsen , Asbjørn Munk

Although electrocardiograms (ECG) play a dominant role in cardiovascular diagnosis and treatment, their intrinsic data forms and representational patterns pose significant challenges for medical multimodal large language models (Med-MLLMs)…

Despite that deep learning (DL) methods have presented tremendous potential in many medical image analysis tasks, the practical applications of medical DL models are limited due to the lack of enough data samples with manual annotations. By…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Zhiyang Liu , Dong Yang , Minghao Zhang , Hanyu Sun , Hong Wu , Huiying Wang , Wen Shen , Chao Chai , Shuang Xia

Emotion understanding is an essential but highly challenging component of artificial general intelligence. The absence of extensively annotated datasets has significantly impeded advancements in this field. We present EmotionCLIP, the first…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Sitao Zhang , Yimu Pan , James Z. Wang

Multimodal models excel in English, supported by abundant image-text and audio-text data, but performance drops sharply for other languages due to limited multilingual multimodal resources. Existing solutions rely on machine translation,…

Machine Learning · Computer Science 2026-01-22 Piyush Singh Pasi

Multimodal large language models (MLLMs) are increasingly being applied in the medical field, particularly in medical imaging. However, developing MLLMs for ECG signals, which are crucial in clinical settings, has been a significant…

Computation and Language · Computer Science 2024-11-25 Haitao Li , Ziyu Li , Yiheng Mao , Ziyi Liu , Zhoujian Sun , Zhengxing Huang

Machine learning models built on training data with multiple modalities can reveal new insights that are not accessible through unimodal datasets. For example, cardiac magnetic resonance images (MRIs) and electrocardiograms (ECGs) are both…

Machine Learning · Computer Science 2023-12-05 Bum Chul Kwon , Samuel Friedman , Kai Xu , Steven A Lubitz , Anthony Philippakis , Puneet Batra , Patrick T Ellinor , Kenney Ng

Current artificial intelligence models for medical imaging are predominantly single modality and single disease. Attempts to create multimodal and multi-disease models have resulted in inconsistent clinical accuracy. Furthermore, training…

Early identification of stroke symptoms is essential for enabling timely intervention and improving patient outcomes, particularly in prehospital settings. This study presents a fast, non-invasive multimodal deep learning framework for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Ngoc-Khai Hoang , Thi-Nhu-Mai Nguyen , Huy-Hieu Pham

Image-text matching is a key multimodal task that aims to model the semantic association between images and text as a matching relationship. With the advent of the multimedia information age, image, and text data show explosive growth, and…

Machine Learning · Computer Science 2024-06-24 Jinyin Wang , Haijing Zhang , Yihao Zhong , Yingbin Liang , Rongwei Ji , Yiru Cang

Cardiovascular magnetic resonance imaging is a powerful diagnostic tool for assessing cardiac structure and function. However, traditional breath-held imaging protocols pose challenges for patients with arrhythmias or limited breath-holding…

Foundation models, large-scale, pre-trained deep-learning models adapted to a wide range of downstream tasks have gained significant interest lately in various deep-learning problems undergoing a paradigm shift with the rise of these…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Bobby Azad , Reza Azad , Sania Eskandari , Afshin Bozorgpour , Amirhossein Kazerouni , Islem Rekik , Dorit Merhof

In echocardiography (echo), an electrocardiogram (ECG) is conventionally used to temporally align different cardiac views for assessing critical measurements. However, in emergencies or point-of-care situations, acquiring an ECG is often…

Computer Vision and Pattern Recognition · Computer Science 2021-02-05 Fatemeh Taheri Dezaki , Christina Luong , Tom Ginsberg , Robert Rohling , Ken Gin , Purang Abolmaesumi , Teresa Tsang

Semi-supervised image classification has shown substantial progress in learning from limited labeled data, but recent advances remain largely untested for clinical applications. Motivated by the urgent need to improve timely diagnosis of…

Computer Vision and Pattern Recognition · Computer Science 2021-08-03 Zhe Huang , Gary Long , Benjamin Wessler , Michael C. Hughes

In clinical practice of echocardiography examinations, multiple planes containing the heart structures of different view are usually required in screening, diagnosis and treatment of cardiac disease. AI models for echocardiography have to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Jiongtong Hu , Wei Zhuo , Jun Cheng , Yingying Liu , Wufeng Xue , Dong Ni

Multimodal language modeling has enabled breakthroughs for representation learning, yet remains unexplored in the realm of functional brain data for clinical phenotyping. This paper pioneers EEG-language models (ELMs) trained on clinical…

Signal Processing · Electrical Eng. & Systems 2025-08-12 Sam Gijsen , Kerstin Ritter

Large-scale vision-language models demonstrate strong multimodal alignment and generalization across diverse tasks. Among them, CLIP stands out as one of the most successful approaches. In this work, we extend the application of CLIP to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Sooyoung Park , Arda Senocak , Joon Son Chung

Due to the lack of paired samples and the low signal-to-noise ratio of functional MRI (fMRI) signals, reconstructing perceived natural images or decoding their semantic contents from fMRI data are challenging tasks. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Yulong Liu , Yongqiang Ma , Wei Zhou , Guibo Zhu , Nanning Zheng

Here we present a versatile foundation model that can perform a range of clinically-relevant image analysis tasks, including segmentation, landmark localisation, diagnosis, and prognostication. A multi-view convolution-transformer masked…

Image and Video Processing · Electrical Eng. & Systems 2025-09-03 Yunguan Fu , Wenjia Bai , Weixi Yi , Charlotte Manisty , Anish N Bhuva , Thomas A Treibel , James C Moon , Matthew J Clarkson , Rhodri Huw Davies , Yipeng Hu

In the evolving landscape of ECG signal analysis, the challenge of limited transparency in machine learning models remains a significant barrier to their effective integration into clinical practice. This study addresses this issue by…

Signal Processing · Electrical Eng. & Systems 2024-12-09 Toygar Tanyel , Sezgin Atmaca , Kaan Gökçe , M. Yiğit Balık , Arda Güler , Emre Aslanger , İlkay Öksüz
‹ Prev 1 8 9 10 Next ›