中文
相关论文

相关论文: EchoVLM: Measurement-Grounded Multimodal Learning …

200 篇论文

Mammography is the primary imaging tool for breast cancer diagnosis. Despite significant strides in applying deep learning to interpret mammography images, efforts that focus predominantly on visual features often struggle with…

图像与视频处理 · 电气工程与系统科学 2024-09-25 Xin Wei , Yaling Tao , Changde Du , Gangming Zhao , Yizhou Yu , Jinpeng Li

Recently, histopathology vision-language foundation models (VLMs) have gained popularity due to their enhanced performance and generalizability across different downstream tasks. However, most existing histopathology benchmarks are either…

图像与视频处理 · 电气工程与系统科学 2025-03-18 Roba Al Majzoub , Hashmat Malik , Muzammal Naseer , Zaigham Zaheer , Tariq Mahmood , Salman Khan , Fahad Khan

The prevalence of vision-threatening eye diseases is a significant global burden, with many cases remaining undiagnosed or diagnosed too late for effective treatment. Large vision-language models (LVLMs) have the potential to assist in…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Zhenyue Qin , Yu Yin , Dylan Campbell , Xuansheng Wu , Ke Zou , Yih-Chung Tham , Ninghao Liu , Xiuzhen Zhang , Qingyu Chen

Vision-Language Models (VLMs) trained via contrastive learning have achieved notable success in natural image tasks. However, their application in the medical domain remains limited due to the scarcity of openly accessible, large-scale…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Muhammad Uzair Khattak , Shahina Kunhimon , Muzammal Naseer , Salman Khan , Fahad Shahbaz Khan

Multimodal language models (MLMs) show promise for clinical decision support and diagnostic reasoning, raising the prospect of end-to-end automated medical image interpretation. However, clinicians are highly selective in adopting AI tools;…

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains,…

Reliable interpretation of echocardiography (Echo) is crucial for assessing cardiac function, which demands clinicians to synchronously orchestrate multiple capabilities, including visual observation (eyes), manual measurement (hands), and…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Qin Wang , Zhiqing He , Yu Liu , Bowen Guo , Zeju Li , Miao Zhao , Wenhao Ju , Zhiling Luo , Xianhong Shu , Yi Guo , Yuanyuan Wang

Foundation models are becoming increasingly effective in the medical domain, offering pre-trained models on large datasets that can be readily adapted for downstream tasks. Despite progress, fetal ultrasound images remain a challenging…

Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We posit this limitation arises from the scarcity of high-quality, large-scale clinical…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Mengmeng Zhang , Xiaoping Wu , Hao Luo , Fan Wang , Yisheng Lv

Cardiovascular diseases stand as the primary global cause of mortality. Among the various imaging techniques available for visualising the heart and evaluating its function, echocardiograms emerge as the preferred choice due to their safety…

图像与视频处理 · 电气工程与系统科学 2023-11-22 Adil Dahlan , Cyril Zakka , Abhinav Kumar , Laura Tang , Rohan Shad , Robyn Fong , William Hiesinger

Video-based Clinical Gait Analysis often suffers from poor generalization as models overfit environmental biases instead of capturing pathological motion. To address this, we propose BioGait-VLM, a tri-modal Vision-Language-Biomechanics…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Erdong Chen , Yuyang Ji , Jacob K. Greenberg , Benjamin Steel , Faraz Arkam , Abigail Lewis , Pranay Singh , Feng Liu

The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-assisted reporting. While existing vision-language models (VLMs) show promise in specific tasks…

Facial video-based remote physiological measurement is a promising research area for detecting human vital signs (e.g., heart rate, respiration frequency) in a non-contact way. Conventional approaches are mostly supervised learning,…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Zijie Yue , Miaojing Shi , Hanli Wang , Shuai Ding , Qijun Chen , Shanlin Yang

Vision-and-language models (VLMs) have been increasingly explored in the medical domain, particularly following the success of CLIP in general domain. However, unlike the relatively straightforward pairing of 2D images and text, curating…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Ziyang Zhang , Yang Yu , Xulei Yang , Si Yong Yeo

Recent advances in large language models (LLMs) have enabled the development of multimodal medical AI. While models such as MedGemini achieve high accuracy on VQA tasks like USMLE MM, their performance on ECG based tasks remains limited,…

机器学习 · 计算机科学 2026-02-12 Junichiro Takahashi , Masataka Sato , Satoshi Kodeta , Norihiko Takeda

Cardiac biosignals, such as electrocardiograms (ECG) and photoplethysmograms (PPG), are of paramount importance for the diagnosis, prevention, and management of cardiovascular diseases, and have been extensively used in a variety of…

Echocardiography interpretation requires integrating multi-view temporal evidence with quantitative measurements and guideline-grounded reasoning, yet existing foundation-model pipelines largely solve isolated subtasks and fail when tool…

Electrocardiogram (ECG) is a widely used tool for assessing cardiac function due to its low cost and accessibility. Emergent research shows that ECGs can help make predictions on key outcomes traditionally derived from more complex…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yuan Gao , Sangwook Kim , Chris McIntosh

Electrocardiogram (ECG) monitoring is one of the most powerful technique of cardiovascular disease (CVD) early identification, and the introduction of intelligent wearable ECG devices has enabled daily monitoring. However, due to the need…

信号处理 · 电气工程与系统科学 2024-03-08 Hongxiang Gao , Xingyao Wang , Zhenghua Chen , Min Wu , Jianqing Li , Chengyu Liu

Visual Emotion Comprehension (VEC) aims to infer sentiment polarities or emotion categories from affective cues embedded in images. In recent years, Multimodal Large Language Models (MLLMs) have established a popular paradigm in VEC,…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Daiqing Wu , Dongbao Yang , Can Ma , Yu Zhou