中文
相关论文

相关论文: E3D-GPT: Enhanced 3D Visual Foundation for Medical…

200 篇论文

As multimodal language models advance, their application to 3D scene understanding is a fast-growing frontier, driving the development of 3D Vision-Language Models (VLMs). Current methods show strong dependence on object detectors,…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Anna-Maria Halacheva , Jan-Nico Zaech , Xi Wang , Danda Pani Paudel , Luc Van Gool

This paper introduces an innovative methodology for producing high-quality 3D lung CT images guided by textual information. While diffusion-based generative models are increasingly used in medical imaging, current state-of-the-art…

图像与视频处理 · 电气工程与系统科学 2024-10-16 Yanwu Xu , Li Sun , Wei Peng , Shuyue Jia , Katelyn Morrison , Adam Perer , Afrooz Zandifar , Shyam Visweswaran , Motahhare Eslami , Kayhan Batmanghelich

In the pursuit of efficient automated content creation, procedural generation, leveraging modifiable parameters and rule-based systems, emerges as a promising approach. Nonetheless, it could be a demanding endeavor, given its intricate…

计算机视觉与模式识别 · 计算机科学 2024-05-30 Chunyi Sun , Junlin Han , Weijian Deng , Xinlong Wang , Zishan Qin , Stephen Gould

A visual-language model (VLM) pre-trained on natural images and text pairs poses a significant barrier when applied to medical contexts due to domain shift. Yet, adapting or fine-tuning these VLMs for medical use presents considerable…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Aisha Urooj Khan , John Garrett , Tyler Bradshaw , Lonie Salkowski , Jiwoong Jason Jeong , Amara Tariq , Imon Banerjee

Precise spatial understanding from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs), as their visual representations are predominantly semantic and lack explicit geometric grounding. While…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Chanyoung Gwak , Yoonwoo Jeong , Byungwoo Jeon , Hyunseok Lee , Jinwoo Shin , Minsu Cho

With the advent of Vision-Language Models (VLMs), medical artificial intelligence (AI) has experienced significant technological progress and paradigm shifts. This survey provides an extensive review of recent advancements in Medical…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Beria Chingnabe Kalpelbe , Angel Gabriel Adaambiik , Wei Peng

Multi-modal large language models (MLLMs) have been given free rein to explore exciting medical applications with a primary focus on radiology report generation. Nevertheless, the preliminary success in 2D radiology captioning is…

Vision-language models (VLMs) show strong potential for complex diagnostic tasks in medical imaging. However, applying VLMs to multi-organ medical imaging introduces two principal challenges: (1) modality-specific vision-language alignment…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Haowen Zhu , Ning Yin , Xiaogen Zhou

The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundation models, which are also critical elements of multi-modal…

计算机视觉与模式识别 · 计算机科学 2024-01-18 Zhe Chen , Jiannan Wu , Wenhai Wang , Weijie Su , Guo Chen , Sen Xing , Muyan Zhong , Qinglong Zhang , Xizhou Zhu , Lewei Lu , Bin Li , Ping Luo , Tong Lu , Yu Qiao , Jifeng Dai

In medical-data driven learning, 3D convolutional neural networks (CNNs) have started to show superior performance to 2D CNNs in numerous deep learning tasks, proving the added value of 3D spatial information in feature representation.…

图像与视频处理 · 电气工程与系统科学 2024-01-24 Xin Wang , Ruisheng Su , Weiyi Xie , Wenjin Wang , Yi Xu , Ritse Mann , Jungong Han , Tao Tan

Decoding visual information from electroencephalography (EEG) has recently achieved promising results, primarily focusing on reconstructing two-dimensional (2D) images from brain activity. However, the reconstruction of three-dimensional…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Emanuele Balloni , Emanuele Frontoni , Chiara Matti , Marina Paolanti , Roberto Pierdicca , Emiliano Santarnecchi

Advances in GPT-based large language models (LLMs) are revolutionizing natural language processing, exponentially increasing its use across various domains. Incorporating uni-directional attention, these autoregressive LLMs can generate…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Lalithkumar Seenivasan , Mobarakol Islam , Gokul Kannan , Hongliang Ren

This review provides a systematic analysis of comprehensive survey of 3D object detection with vision-language models(VLMs) , a rapidly advancing area at the intersection of 3D vision and multimodal AI. By examining over 100 research…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Ranjan Sapkota , Konstantinos I Roumeliotis , Rahul Harsha Cheppally , Marco Flores Calero , Manoj Karkee

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains challenging and untouched so far. Fortunately, human motion…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Biao Jiang , Xin Chen , Wen Liu , Jingyi Yu , Gang Yu , Tao Chen

Deep learning has been widely used in medical image segmentation and other aspects. However, the performance of existing medical image segmentation models has been limited by the challenge of obtaining sufficient high-quality labeled data…

计算机视觉与模式识别 · 计算机科学 2023-06-28 Zihan Li , Yunxiang Li , Qingde Li , Puyang Wang , Dazhou Guo , Le Lu , Dakai Jin , You Zhang , Qingqi Hong

While emerging 3D medical foundation models are envisioned as versatile tools with offer general-purpose capabilities, their validation remains largely confined to regional and structural imaging, leaving a significant modality discrepancy…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Yichi Zhang , Feiyang Xiao , Le Xue , Wenbo Zhang , Gang Feng , Chenguang Zheng , Yuan Qi , Yuan Cheng , Zixin Hu

This article describes the development of a platform designed to visualize the 3D motion of the tongue using ultrasound image sequences. An overview of the system design is given and promising results are presented. Compared to the analysis…

计算机视觉与模式识别 · 计算机科学 2016-05-20 Kele Xu , Yin Yang , Aurore Jaumard-Hakoun , Clemence Leboullenger , Gerard Dreyfus , Pierre Roussel , Maureen Stone , Bruce Denby

Current medical vision-language models (VLMs) process volumetric brain MRI using 2D slice-based approximations, fragmenting the spatial context required for accurate neuroradiological interpretation. We developed \textbf{Brain3D}, a staged…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Mariano Barone , Francesco Di Serio , Giuseppe Riccio , Antonio Romano , Marco Postiglione , Antonino Ferraro , Vincenzo Moscato

Large language models (LLMs) have recently demonstrated their potential in clinical applications, providing valuable medical knowledge and advice. For example, a large dialog LLM like ChatGPT has successfully passed part of the US medical…

计算机视觉与模式识别 · 计算机科学 2023-02-15 Sheng Wang , Zihao Zhao , Xi Ouyang , Qian Wang , Dinggang Shen

Conditional medical image generation plays an important role in many clinically relevant imaging tasks. However, existing methods still face a fundamental challenge in balancing inference efficiency, patient-specific fidelity, and…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Zirong Li , Siyuan Mei , Weiwen Wu , Andreas Maier , Lina Gölz , Yan Xia