English

EchoAgent: Guideline-Centric Reasoning Agent for Echocardiography Measurement and Interpretation

Computer Vision and Pattern Recognition 2025-11-19 v1 Computation and Language Machine Learning

Abstract

Purpose: Echocardiographic interpretation requires video-level reasoning and guideline-based measurement analysis, which current deep learning models for cardiac ultrasound do not support. We present EchoAgent, a framework that enables structured, interpretable automation for this domain. Methods: EchoAgent orchestrates specialized vision tools under Large Language Model (LLM) control to perform temporal localization, spatial measurement, and clinical interpretation. A key contribution is a measurement-feasibility prediction model that determines whether anatomical structures are reliably measurable in each frame, enabling autonomous tool selection. We curated a benchmark of diverse, clinically validated video-query pairs for evaluation. Results: EchoAgent achieves accurate, interpretable results despite added complexity of spatiotemporal video analysis. Outputs are grounded in visual evidence and clinical guidelines, supporting transparency and traceability. Conclusion: This work demonstrates the feasibility of agentic, guideline-aligned reasoning for echocardiographic video analysis, enabled by task-specific tools and full video-level automation. EchoAgent sets a new direction for trustworthy AI in cardiac ultrasound.

Keywords

Cite

@article{arxiv.2511.13948,
  title  = {EchoAgent: Guideline-Centric Reasoning Agent for Echocardiography Measurement and Interpretation},
  author = {Matin Daghyani and Lyuyang Wang and Nima Hashemi and Bassant Medhat and Baraa Abdelsamad and Eros Rojas Velez and XiaoXiao Li and Michael Y. C. Tsang and Christina Luong and Teresa S. M. Tsang and Purang Abolmaesumi},
  journal= {arXiv preprint arXiv:2511.13948},
  year   = {2025}
}

Comments

12 pages, Under Review