基于多站位听诊记录的患者水平多模态问答
声音
2026-03-17 v1 人工智能
音频与语音处理
摘要
听诊是诊断工具,常受主观解释限制。虽然通用音频语言模型在通用领域表现出色,但在生理信号细微差别上难以胜任。我们提出一种框架,通过门控注意力将多站位听诊记录直接对齐到冻结的大语言模型嵌入空间中。通过利用LLM的潜在世界知识,Our approach moves beyond isolated classification toward holistic, patient-level assessment. On the CaReSound benchmark, our model achieves a state-of-the-art 0.865 F1-macro and 0.952 BERTScore. We demonstrate that lightweight, domain-specific encoders rival large-scale ALMs and that multi-site aggregation provides spatial redundancy that mitigates temporal truncation. This alignment of medical acoustics with text foundations offers a scalable path for bridging signal processing and clinical assessment.
引用
@article{arxiv.2603.13362,
title = {Patient-Level Multimodal Question Answering from Multi-Site Auscultation Recordings},
author = {Fan Wu and Tsai-Ning Wang and Nicolas Zumarraga and Ning Wang and Markus Kreft and Kevin O'Sullivan and Elgar Fleisch and Oliver Aalami and Paul Schmiedmayer and Robert Jakob and Patrick Langer},
journal= {arXiv preprint arXiv:2603.13362},
year = {2026}
}