中文
相关论文

相关论文: Towards Reliable Fetal Ultrasound Interpretation w…

200 篇论文

Fetal ultrasound (US) is the primary imaging modality for prenatal screening, yet its interpretation relies heavily on the expertise of the clinician. Despite advances in deep learning and foundation models, existing automated tools for…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Xiaotian Hu , Junwei Huang , Mingxuan Liu , Kasidit Anmahapong , Yifei Chen , Yitong Luo , Yiming Huang , Xuguang Bai , Zihan Li , Yi Liao , Haibo Qu , Qiyuan Tian

The growing demand for prenatal ultrasound imaging has intensified a global shortage of trained sonographers, creating barriers to essential fetal health monitoring. Deep learning has the potential to enhance sonographers' efficiency and…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Hussain Alasmawi , Numan Saeed , Mohammad Yaqub

Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are adapted to structured adult imaging and underperform in fetal ultrasound, which poses challenges of…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Xiao He , Huangxuan Zhao , Guojia Wan , Wei Zhou , Yanxing Liu , Juhua Liu , Yongchao Xu , Yong Luo , Dacheng Tao , Bo Du

Focused Ultrasound Ablation Surgery (FUAS) has emerged as a promising non-invasive therapeutic modality, valued for its safety and precision. Nevertheless, its clinical implementation entails intricate tasks such as multimodal image…

多智能体系统 · 计算机科学 2026-03-31 Lina Zhao , Zihao Bian , Qingyue Chen , Yafang Li , Zhiyi Luo , Jiaxing Bai , Guangbo Li , Min He , Kezhi Li , Huaiyuan Yao , Zongjiu Zhang

Finite Element Analysis (FEA) serves as the cornerstone of modern engineering design. However, its workflow is inherently complex and relies heavily on domain expertise. Although recent efforts have integrated Large Language Models (LLMs)…

人工智能 · 计算机科学 2026-05-29 Jiachen Zhang , Junyi Lao , Chenghao Liu , Siyuan Liu , Shixin Wu , Linsen Zhang , Boyu Wang , Songfang Huang

Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of these capabilities: specialized detectors offer strong localization but limited…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Jing Zhang , Wentao Jiang , Tao Huang , Zhiwei Wang , Jianxin Liu , Jian Chen , Ping Ye , Gang Wang , Zengmao Wang , Bo Du , Dacheng Tao

Vision-Language Models (VLMs) show promise in medical diagnosis, yet suffer from reasoning detachment, where linguistically fluent explanations drift from verifiable image evidence, undermining clinical trust. Recent multi-agent frameworks…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Qianhan Feng , Zhongzhen Huang , Yakun Zhu , Xiaofan Zhang , Qi Dou

Ultrasound offers a safe, cost-effective, and widely accessible technology for fetal brain imaging, making it especially suitable for routine clinical use. However, it suffers from view-dependent artifacts, operator variability, and a…

图像与视频处理 · 电气工程与系统科学 2026-01-13 Mohammad Khateri , Morteza Ghahremani , Sergio Valencia , Camilo Jaimes , Alejandra Sierra , Jussi Tohka , P. Ellen Grant , Davood Karimi

Purpose: Echocardiographic interpretation requires video-level reasoning and guideline-based measurement analysis, which current deep learning models for cardiac ultrasound do not support. We present EchoAgent, a framework that enables…

Large language models (LLMs) show promise for healthcare question answering, but clinical use is limited by weak verification, insufficient evidence grounding, and unreliable confidence signalling. We propose a multi-agent medical QA…

计算与语言 · 计算机科学 2026-02-17 Naeimeh Nourmohammadi , Md Meem Hossain , The Anh Han , Safina Showkat Ara , Zia Ush Shamszaman

Ultrasound is a cornerstone of emergency and hepatobiliary imaging, yet its interpretation remains highly operator-dependent and time-sensitive. Here, we present a multitask vision-language agent (VLM) developed to assist with comprehensive…

Multimodal Large Language Model (MLLM) has recently garnered attention as a prominent research focus. By harnessing powerful LLM, it facilitates a transition of conversational generative AI from unimodal text to performing multimodal tasks.…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Xuechen Guo , Wenhao Chai , Shi-Yan Li , Gaoang Wang

Multimodal large language models (MLLMs) excel at generating highly detailed captions but often produce hallucinations. Our analysis reveals that existing hallucination detection methods struggle with detailed captions. We attribute this to…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Saehyung Lee , Seunghyun Yoon , Trung Bui , Jing Shi , Sungroh Yoon

Real-world multimodal applications often require any-to-any capabilities, enabling both understanding and generation across modalities including text, image, audio, and video. However, integrating the strengths of autoregressive language…

机器学习 · 计算机科学 2025-08-15 Jiulin Li , Ping Huang , Yexin Li , Shuo Chen , Juewen Hu , Ye Tian

Document Question Answering (DocQA) is a very common task. Existing methods using Large Language Models (LLMs) or Large Vision Language Models (LVLMs) and Retrieval Augmented Generation (RAG) often prioritize information from a single…

机器学习 · 计算机科学 2025-03-19 Siwei Han , Peng Xia , Ruiyi Zhang , Tong Sun , Yun Li , Hongtu Zhu , Huaxiu Yao

3D CT analysis spans a continuum from low-level perception to high-level clinical understanding. Existing 3D-oriented analysis methods adopt either isolated task-specific modeling or task-agnostic end-to-end paradigms to produce one-hop…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Ziyue Wang , Linghan Cai , Chang Han Low , Haofeng Liu , Junde Wu , Jingyu Wang , Rui Wang , Lei Song , Jiang Bian , Jingjing Fu , Yueming Jin

The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologist shortages and improve clinical reporting efficiency. However, applying general LLMs to…

多智能体系统 · 计算机科学 2025-12-18 Philip R. Liu , Sparsh Bansal , Jimmy Dinh , Aditya Pawar , Ramani Satishkumar , Shail Desai , Neeraj Gupta , Xin Wang , Shu Hu

In this paper, we propose an end-to-end multi-task neural network called FetalNet with an attention mechanism and stacked module for spatio-temporal fetal ultrasound scan video analysis. Fetal biometric measurement is a standard examination…

图像与视频处理 · 电气工程与系统科学 2022-05-04 Szymon Płotka , Tomasz Włodarczyk , Adam Klasa , Michał Lipa , Arkadiusz Sitek , Tomasz Trzciński

Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, existing medical visual question answering (VQA) benchmarks collapse model capabilities…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Yixiong Chen , Wenjie Xiao , Pedro R. A. S. Bassi , Boyan Wang , Liang He , Xinze Zhou , Sezgin Er , Ibrahim Ethem Hamamci , Zongwei Zhou , Alan Yuille

Large language models (LLMs) have demonstrated immense capabilities in understanding textual data and are increasingly being adopted to help researchers accelerate scientific discovery through knowledge extraction (information retrieval),…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Robinson Umeike , Neil Getty , Fangfang Xia , Rick Stevens
‹ 上一页 1 2 3 10 下一页 ›