English
Related papers

Related papers: TotalFM: An Organ-Separated Framework for 3D-CT Vi…

200 papers

Accurate delineation of anatomical structures in volumetric CT scans is crucial for diagnosis and treatment planning. While AI has advanced automated segmentation, current approaches typically target individual structures, creating a…

Foundation models (FMs) and large language models (LLMs) have demonstrated promising generalization across diverse domains for time-series analysis, yet their potential for electronic fetal monitoring (EFM) and cardiotocography (CTG)…

Machine Learning · Computer Science 2025-11-07 Sheng Wong , Ravi Shankar , Beth Albert , Gabriel Davis Jones

Learning by imitation is one of the most significant abilities of human beings and plays a vital role in human's computational neural system. In medical image analysis, given several exemplars (anchors), experienced radiologist has the…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Hong-Yu Zhou , Hualuo Liu , Shilei Cao , Dong Wei , Chixiang Lu , Yizhou Yu , Kai Ma , Yefeng Zheng

Recent advances in 3D fully convolutional networks (FCN) have made it feasible to produce dense voxel-wise predictions of volumetric images. In this work, we show that a multi-class 3D FCN trained on manually labeled CT scans of several…

Computer Vision and Pattern Recognition · Computer Science 2018-04-13 Holger R. Roth , Hirohisa Oda , Xiangrong Zhou , Natsuki Shimizu , Ying Yang , Yuichiro Hayashi , Masahiro Oda , Michitaka Fujiwara , Kazunari Misawa , Kensaku Mori

Foundation models for image segmentation have shown strong generalization in natural images, yet their applicability to 3D medical imaging remains limited. In this work, we study the zero-shot use of Segment Anything Model 2 (SAM2) for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Miquel Lopez Escoriza , Pau Amargant Alvarez

Foundation models for medical imaging demonstrate superior generalization capabilities across diverse anatomical structures and clinical applications. Their outstanding performance relies on substantial computational resources, limiting…

Image and Video Processing · Electrical Eng. & Systems 2026-04-15 Chen Ma , Jing Jiao , Shuyu Liang , Junhu Fu , Qin Wang , Zeju Li , Yuanyuan Wang , Yi Guo

Computed tomography (CT) and clinical numeric data are essential modalities for cancer evaluation, but building large-scale multimodal training datasets for developing medical foundation models remains challenging due to the structural…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Daeun Jung , Jaehyeok Jang , Sooyoung Jang , Yu Rang Park

Current artificial intelligence models for medical imaging are predominantly single modality and single disease. Attempts to create multimodal and multi-disease models have resulted in inconsistent clinical accuracy. Furthermore, training…

Vision foundation models have achieved remarkable progress across various image analysis tasks. In the image segmentation task, foundation models like the Segment Anything Model (SAM) enable generalizable zero-shot segmentation through…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Xingxin He , Yifan Hu , Zhaoye Zhou , Mohamed Jarraya , Fang Liu

Foundation models or pre-trained models have substantially improved the performance of various language, vision, and vision-language understanding tasks. However, existing foundation models can only perform the best in one type of tasks,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Xinsong Zhang , Yan Zeng , Jipeng Zhang , Hang Li

Ultrasound imaging is generally employed for real-time investigation of internal anatomy of the human body for disease identification. Delineation of the anatomical boundary of organs and pathological lesions is quite challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2019-01-23 Sumanth Nandamuri , Debarghya China , Pabitra Mitra , Debdoot Sheet

Vision-language models (VLMs) have shown promise in 2D medical image analysis, but extending them to 3D remains challenging due to the high computational demands of volumetric data and the difficulty of aligning 3D spatial features with…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Yu Xin , Gorkem Can Ates , Kuang Gong , Wei Shao

Multimodal Large Language Models (MLLMs) have emerged as a promising way to automate Radiology Report Generation (RRG). In this work, we systematically investigate the design space of 3D MLLMs, including visual input representation,…

Image and Video Processing · Electrical Eng. & Systems 2025-09-23 Mohammed Baharoon , Jun Ma , Congyu Fang , Augustin Toma , Bo Wang

Traditional radio map estimation (RME) techniques fail to capture multi-dimensional and dynamic characteristics of complex spectrum environments. Recent data-driven methods achieve accurate RME in spatial domain, but ignore physical prior…

Signal Processing · Electrical Eng. & Systems 2026-02-27 Dong Yang , Yue Wang , Songyang Zhang , Yingshu Li , Zhipeng Cai , Zhi Tian

Foundation models, trained on vast amounts of data using self-supervised techniques, have emerged as a promising frontier for advancing artificial intelligence (AI) applications in medicine. This study evaluates three different…

Multi-modal medical imaging enables comprehensive diagnostics, yet current foundation models process 2D (e.g. X-ray) and 3D (e.g. CT) data with separate, dimensionality-specific architectures. We present MultiMedVision, a unified framework…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Frank Li , Bardia Khosravi , Mohammadreza Chavoshi , Young Seok Jeon , Theo Dapamede , Hari Trivedi , Janice Newsome , Judy Gichoya

The Segment Anything Model (SAM) and CLIP are remarkable vision foundation models (VFMs). SAM, a prompt driven segmentation model, excels in segmentation tasks across diverse domains, while CLIP is renowned for its zero shot recognition…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Sidra Aleem , Fangyijie Wang , Mayug Maniparambil , Eric Arazo , Julia Dietlmeier , Guenole Silvestre , Kathleen Curran , Noel E. O'Connor , Suzanne Little

Building a large-scale training dataset is an essential problem in the development of medical image recognition systems. Visual grounding techniques, which automatically associate objects in images with corresponding descriptions, can…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Akimichi Ichinose , Taro Hatsutani , Keigo Nakamura , Yoshiro Kitamura , Satoshi Iizuka , Edgar Simo-Serra , Shoji Kido , Noriyuki Tomiyama

A Multistage Full Matching disparity estimation scheme (MFM) is proposed in this work. We demonstrate that decouple all similarity scores directly from the low-resolution 4D volume step by step instead of estimating low-resolution 3D cost…

Computer Vision and Pattern Recognition · Computer Science 2020-12-11 Hong Zhang , Shenglun Chen , Zhihui Wang , Haojie Li , Wanli Ouyang

Medical image segmentation plays a crucial role in AI-assisted diagnostics, surgical planning, and treatment monitoring. Accurate and robust segmentation models are essential for enabling reliable, data-driven clinical decision making…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Sachin Dudda Nagaraju , Ashkan Moradi , Bendik Skarre Abrahamsen , Mattijs Elschot