中文
相关论文

相关论文: OmniCT: Towards a Unified Slice-Volume LVLM for Co…

200 篇论文

Efficient and accurate multi-organ segmentation from abdominal CT volumes is a fundamental challenge in medical image analysis. Existing 3D segmentation approaches are computationally and memory intensive, often processing entire volumes…

图像与视频处理 · 电气工程与系统科学 2025-05-19 Hania Ghouse , Muzammil Behzad

The practical deployment of medical vision-language models (Med-VLMs) necessitates seamless integration of textual data with diverse visual modalities, including 2D/3D images and videos, yet existing models typically employ separate…

计算与语言 · 计算机科学 2025-04-22 Songtao Jiang , Yuan Wang , Sibo Song , Yan Zhang , Zijie Meng , Bohan Lei , Jian Wu , Jimeng Sun , Zuozhu Liu

Accurate disease interpretation from radiology remains challenging due to imaging heterogeneity. Achieving expert-level diagnostic decisions requires integration of subtle image features with clinical knowledge. Yet major vision-language…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Difei Gu , Yunhe Gao , Mu Zhou , Dimitris Metaxas

Interpreting quantitative CT biomarkers, such as organ volume and tissue attenuation, requires large-scale healthy reference distributions. However, creating these is challenging because clinical datasets are often heavily enriched with…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Christian Wachinger , Bernhard Renger , Christopher Späth , Jan Kirschke , Marcus Makowski

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in various multimodal tasks. However, their potential in the medical domain remains largely unexplored. A significant challenge arises from the scarcity of…

图像与视频处理 · 电气工程与系统科学 2024-04-23 Yutao Hu , Tianbin Li , Quanfeng Lu , Wenqi Shao , Junjun He , Yu Qiao , Ping Luo

Recent medical vision-language models (VLMs) have shown promise in 2D medical image interpretation. However extending them to 3D medical imaging has been challenging due to computational complexities and data scarcity. Although a few recent…

图像与视频处理 · 电气工程与系统科学 2024-12-19 Changsun Lee , Sangjoon Park , Cheong-Il Shin , Woo Hee Choi , Hyun Jeong Park , Jeong Eun Lee , Jong Chul Ye

General-purpose vision-language models (VLMs) have emerged as promising tools in radiology, offering zero-shot capabilities that mitigate the need for large labeled datasets. However, in high-stakes domains like diagnostic radiology, these…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Hao-Chih Lee , Zelong Liu , Hamza Ahmed , Spencer Kim , Sean Huver , Vishwesh Nath , Zahi A. Fayad , Timothy Deyer , Xueyan Mei

Magnetic Resonance Imaging (MRI) is indispensable in clinical practice but remains constrained by fragmented, multi-stage workflows encompassing acquisition, reconstruction, segmentation, detection, diagnosis, and reporting. While deep…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Xingxin He , Aurora Rofena , Ruimin Feng , Haozhe Liao , Zhaoye Zhou , Albert Jang , Fang Liu

Large-scale, volumetric medical imaging datasets typically aggregate scans from different vendors and devices, resulting in highly variable resolution, slice thicknesses, and numbers of slices per study. Consequently, training…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Jiayi Wang , Hadrien Reynaud , Ibrahim Ethem Hamamci , Sezgin Er , Suprosanna Shit , Bjoern Menze , Bernhard Kainz

Visual question answering (VQA) in medical imaging aims to support clinical diagnosis by automatically interpreting complex imaging data in response to natural language queries. Existing studies typically rely on distinct visual and textual…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Yuanhe Tian , Chen Su , Junwen Duan , Yan Song

Computed tomography (CT) is extensively used for accurate visualization and segmentation of organs and lesions. While deep learning models such as convolutional neural networks (CNNs) and vision transformers (ViTs) have significantly…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Yuheng Li , Yuxiang Lai , Maria Thor , Deborah Marshall , Zachary Buchwald , David S. Yu , Xiaofeng Yang

Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Yang Xing , Jiong Wu , Savas Ozdemir , Ying Zhang , Yang Yang , Wei Shao , Kuang Gong

Computed Tomography (CT) is one of the most popular modalities for medical imaging. By far, CT images have contributed to the largest publicly available datasets for volumetric medical segmentation tasks, covering full-body anatomical…

图像与视频处理 · 电气工程与系统科学 2024-11-25 Jin Ye , Ying Chen , Yanjun Li , Haoyu Wang , Zhongying Deng , Ziyan Huang , Yanzhou Su , Chenglong Ma , Yuanfeng Ji , Junjun He

Recent advances in 3D medical vision-language models have enabled joint reasoning over volumetric images and text, showing strong performance in medical visual question-answering (VQA) and report generation. Despite this progress, it…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Mashrafi Monon , Umaima Rahman , Asif Hanif , Numan Saeed , Mohammad Yaqub

With the rapid advancement of deep learning, particularly in the field of medical image analysis, an increasing number of Vision-Language Models (VLMs) are being widely applied to solve complex health and biomedical challenges. However,…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Haiyang Yu , Siyang Yi , Ke Niu , Minghan Zhuo , Bin Li

Recent advances in multimodal large language models (LLMs) have highlighted their potential for medical and surgical applications. However, existing surgical datasets predominantly adopt a Visual Question Answering (VQA) format with…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Tae-Min Choi , Tae Kyeong Jeong , Garam Kim , Jaemin Lee , Yeongyoon Koh , In Cheul Choi , Jae-Ho Chung , Jong Woong Park , Juyoun Park

The advancement of artificial intelligence (AI) for organ segmentation and tumor detection is propelled by the growing availability of computed tomography (CT) datasets with detailed, per-voxel annotations. However, these AI models often…

图像与视频处理 · 电气工程与系统科学 2024-05-29 Jie Liu , Yixiao Zhang , Kang Wang , Mehmet Can Yavuz , Xiaoxi Chen , Yixuan Yuan , Haoliang Li , Yang Yang , Alan Yuille , Yucheng Tang , Zongwei Zhou

Surgical scene understanding is a cornerstone of computer-assisted intervention. While recent advances, particularly in surgical image segmentation, have driven progress, real-world clinical applications require a more holistic…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Jincai Huang , Shihao Zou , Yuchen Guo , Jingjing Li , Wei Ji , Kai Wang , Shanshan Wang , Weixin Si

Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study…

‹ 上一页 1 2 3 10 下一页 ›